Information processing device, information processing system, program, and information processing method

WO2026181331A1PCT designated stage Publication Date: 2026-09-03MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021053
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-25
Filing Date
2025-06-11
Publication Date
2026-09-03

Smart Images

  • Figure JP2025021053_03092026_PF_FP_ABST
    Figure JP2025021053_03092026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device (100) comprises: an image information acquisition unit (103) that acquires image information indicating an image; a patch information acquisition unit (104) that acquires first patch information indicating a first patch image, and second patch information indicating a second patch image when the image is divided into a plurality of patch images including the first patch image and the second patch image; an image conversion unit (105) that converts each of the image information, the first patch information, and the second patch information into a feature amount; an importance determination unit (111) that determines the importance of each patch image on the basis of the first patch feature amount and the second patch feature amount; and a feature amount extraction unit (112) that extracts a feature amount corresponding to a specific region of the image from the image feature amount, on the basis of the determination result by the importance determination unit.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing System, Program, and Information Processing Method

[0001] The present disclosure relates to an information processing apparatus, an information processing system, a program, and an information processing method.

[0002] Conventionally, a moving object control apparatus for specifying a stopping position of an electric kickboard using a trained model that has appropriately learned the correspondence between language and images has been disclosed (see, for example, Patent Document 1). The trained model used in this moving object control apparatus is a trained model including a pre-trained visual language model that is trained so as to output, when a captured image obtained by capturing an image of the surroundings of the moving object control apparatus with a camera mounted on the moving object control apparatus and an input instruction sentence input by a user of the moving object control apparatus are input, the stopping position of the moving object control apparatus corresponding to the instruction sentence in the image.

[0003] Japanese Unexamined Patent Application Publication No. 2024-031978

[0004] In general, in a visual language model that generates information related to image information and language information based on inputs of the image information and the language information, when the image information includes a plurality of pieces of information indicating objects, people, characters, backgrounds, and the like, there is a problem that hallucination, in which the visual language model generates information different from information desired by a user of the visual language model, may occur.

[0005] The present disclosure has been made based on recognition of the above problem, and an object of the present disclosure is to provide an information processing apparatus, an information processing system, a program, and an information processing method that can suppress hallucination when processing image information using a trained model more effectively than conventional techniques.

[0006] The information processing device according to this disclosure is characterized by comprising: an image information acquisition unit that acquires image information representing an image; a patch information acquisition unit that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit that converts the image information, the first patch information and the second patch information into image features, first patch features and second patch features, respectively; an importance determination unit that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; and a feature extraction unit that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit.

[0007] According to this disclosure, hallucination can be suppressed more effectively when processing image information with a pre-trained model than in conventional methods.

[0008] This is a block diagram showing the schematic configuration of the information processing system according to Embodiment 1. This is a block diagram showing an example of the hardware configuration of the information processing device according to Embodiment 1. This is a block diagram showing an example of the hardware configuration of the information processing device according to Embodiment 1. This is a flowchart showing an example of the processing performed by the information processing device according to Embodiment 1. This is a diagram showing an example of the original image related to the image information acquired by the information processing device according to Embodiment 1. This is a diagram showing an example of the original image related to the image information acquired by the information processing device according to Embodiment 1 being divided. This is a diagram showing an example of the state in which the highly important regions of the original image related to the image information acquired by the information processing device according to Embodiment 1 have been extracted. This is a block diagram showing the schematic configuration of the information processing system according to Embodiment 2. This is a flowchart showing an example of the processing performed by the information processing device according to Embodiment 2. This is a diagram showing an example of the results of morphological analysis by the information processing device according to Embodiment 2. This is a diagram showing an example of the state in which the highly important regions of the original image related to the image information acquired by the information processing device according to Embodiment 2 have been extracted. This is a block diagram showing the schematic configuration of the information processing systems according to Embodiments 3 and 4. This is a flowchart showing an example of the processing performed by the information processing device according to Embodiment 3. Figure 14A shows an example of the similarity between subject information and each patch information acquired by the information processing device according to Embodiment 3, and Figure 14B shows an example of a state in which some patch information has been extracted from multiple patch information acquired by the information processing device according to Embodiment 3 based on the similarity with subject information. This is a flowchart of an example of processing performed by the information processing device according to Embodiment 4. This is a block diagram showing the schematic configuration of the information processing system according to Embodiment 5. This is a flowchart of an example of processing performed by the information processing device according to Embodiment 5.

[0009] Hereinafter, embodiments relating to this disclosure will be described in detail with reference to the drawings. Embodiment 1. First, the information processing system 1 according to Embodiment 1 will be described with reference to Figure 1. Figure 1 is a block diagram showing the schematic configuration of the information processing system 1 according to Embodiment 1. The information processing system according to Embodiment 1 is a system for generating information relating to a specific image based on information indicated in natural language input by a user. As shown in Figure 1, the information processing system 1 according to Embodiment 1 is equipped with an input device 10 for inputting information, a display device 20 for displaying information, a model processing device 30, and an information processing device 100, and these are connected wirelessly or by wire to enable communication with each other. The input device 10, the display device 20, the model processing device 30, and the information processing device 100 may be connected to each other so as to enable communication of information via other devices or communication lines not shown.

[0010] The input device 10 inputs various types of information to the display device 20, the model processing device 30, and the information processing device 100. For example, the input device 10 is composed of one or more combinations of devices such as a keyboard, mouse, touch panel, mechanical switch, camera, microphone, and storage device for storing information, and inputs various types of information to the display device 20, the model processing device 30, and the information processing device 100. For example, the input device 10 receives an input operation from the user and inputs string information indicated by a string of natural language corresponding to the input operation, and image information indicating an image, to the display device 20, the model processing device 30, and the information processing device 100. Also, for example, the input device 10 inputs string information and image information acquired by the functions of the input device 10 to the display device 20, the model processing device 30, and the information processing device 100. Specifically, the input device 10 inputs request information, which is string information indicating a request (prompt) from the user in natural language, and image information indicating an image, to the model processing device 30. In Embodiment 1, "image" includes still images and moving images.

[0011] The display device 20 displays various types of information input from the input device 10, the model processing device 30, and the information processing device 100. In other words, the display device 20 outputs various types of information input from the input device 10, the model processing device 30, and the information processing device 100 as visual information. For example, the display device 20 is composed of a liquid crystal display panel, an organic or inorganic EL (Electroluminescence) panel, a dot matrix display, an LED (Light Emitting Diode), etc., and displays information indicated by natural language strings input from the input device 10, the model processing device 30, and the information processing device 100, as well as image information indicating images. If the input device 10 is composed of a touch panel, the display device 20 may be configured integrally with the input device 10 as a liquid crystal display panel, etc.

[0012] The model processing unit 30 has an information conversion unit 31 and a trained model 32, and generates information corresponding to the input information based on the information input from the input device 10 and the information processing unit 100, and outputs the generated information toward the display device 20 or the information processing unit 100. For example, the model processing unit 30 is composed of a cloud server, a physical server, or other computer. The model processing unit 30 may be configured to acquire information from other devices (not shown) that are connected wirelessly or by wire to enable communication with each other, and may be configured to output the acquired information toward the display device 20 and the information processing unit 100. For example, the model processing unit 30 may be configured to acquire information corresponding to the input information based on the information input from the input device 10 from other devices (not shown), and output the acquired information toward the display device 20 and the information processing unit 100.

[0013] The information conversion unit 31, acting as a feature transformation unit, converts the information input to the model processing unit 30 into information that can be processed by the trained model 32. For example, the information conversion unit 31 is configured as a multilayer perceptron (MLP) and converts the information input to the model processing unit 30 into information that can be processed by the trained model 32. For example, the information conversion unit 31 converts information indicating image features, input from the information processing unit 100, into information that can be processed by the trained model 32. Details regarding the information indicating image features input from the information processing unit 100 to the information conversion unit 31 will be described later. Furthermore, the information conversion unit 31 only needs to be configured to convert the information input to the model processing unit 30 into information that can be processed by the trained model 32, and may be configured as an architecture or algorithm other than a multilayer perceptron. For example, the information conversion unit 31 may be configured to convert the data into information that can be processed by the trained model 32, using a transformer, convolutional neural network, or the like, which has learned the relationships between different input data.

[0014] The trained model 32 receives input string information expressed in natural language and generates information corresponding to the input information. For example, the trained model 32 is composed of Large Language Models. For example, the trained model 32 receives input string information expressed in natural language from the input device 10 and the information processing device 100 and generates information corresponding to the input information. Also, for example, the trained model 32 receives input information from the information conversion unit 31 and generates information corresponding to the input information.

[0015] Furthermore, the information conversion unit 31 may also function as a vision encoder, converting image information input from the input device 10 or other devices (not shown) into feature quantities represented by high-dimensional numerical vectors. The model processing device 30 may be configured to function as a VLM (Vision and Language Model) that accepts input string information and image information expressed in natural language, and generates information corresponding to the input information, through such an information conversion unit 31 and a trained model 32.

[0016] The information processing device 100 includes an input unit 101, an output unit 102, an image information acquisition unit 103, a patch information acquisition unit 104, an image conversion unit 105, an importance determination unit 111, and a feature extraction unit 112.

[0017] The input unit 101 receives various types of information input to the information processing device 100. For example, the input unit 101 receives various types of information input to the information processing device 100 from the input device 10 and the model processing device 30. In addition to the information from the input device 10 and the model processing device 30, the input unit 101 may also be configured to receive information input to the information processing device 100 from other devices (not shown) that are connected to the information processing device 100 wirelessly or by wire so as to be able to communicate with each other.

[0018] The output unit 102 outputs various types of information from the information processing device 100 to the display device 20 and the model processing device 30. For example, the output unit 102 outputs information indicating the results of processing from the information processing device 100 to the display device 20 in order to display the results of processing performed by the information processing device 100 on the display device 20. Also, for example, the output unit 102 outputs information indicating the results of processing from the information processing device 100 to the model processing device 30 in order to cause the trained model 32 to generate information based on the results of processing performed by the information processing device 100. In addition to the display device 20 and the model processing device 30, the output unit 102 may be configured to output information to other devices (not shown) that are wirelessly or wired and connected to the information processing device 100 so that they can communicate with each other.

[0019] The image information acquisition unit 103 acquires image information representing an image via the input unit 101. For example, the image information acquisition unit 103 acquires image information input from the input device 10 via the input unit 101. Also, for example, the image information acquisition unit 103 acquires image information corresponding to the input information from another device (not shown) via the input unit 101, based on the information input from the input device 10. Also, for example, the image information acquisition unit 103 acquires image information input from the model processing device 30 via the input unit 101. The source image related to the image information acquired by the image information acquisition unit 103 is composed of, for example, an object, a person, text, and a landscape, or a combination of several of these.

[0020] The patch information acquisition unit 104 acquires multiple patch information, each representing a patch image when the image corresponding to the image information is divided into multiple patch images, based on the image information acquired by the image information acquisition unit 103. In other words, the patch information acquisition unit 104 acquires multiple patch information, including first patch information representing the first patch image and second patch information representing the second patch image, when the image corresponding to the image information is divided into multiple patch images, including a first patch image and a second patch image, based on the image information acquired by the image information acquisition unit 103. In Embodiment 1, the image represented by the image information acquired by the image information acquisition unit 103 is also called the "original image". For example, the patch information acquisition unit 104 acquires multiple patch information by setting boundaries when the original image is divided into multiple patch images. Alternatively, for example, the patch information acquisition unit 104 acquires multiple patch information by newly generating multiple image information, each representing a patch image when the original image is divided into multiple patch images. The patch information acquisition unit 104 may be configured to output the acquired patch information to an external device of the information processing device 100, such as a display device 20 or a model processing device 30, via an output unit 102.

[0021] The image conversion unit 105 obtains feature quantities representing the original image and the multiple patch images by converting each of the image information acquired by the image information acquisition unit 103 and the multiple patch information acquired by the patch information acquisition unit 104 into image feature quantities, each representing an image feature quantity, a first patch feature quantity, and a second patch feature quantity. In other words, the image conversion unit 105 obtains feature quantities representing the original image and the multiple patch images by converting each of the image information acquired by the image information acquisition unit 103 and the multiple patch information, including the first patch information and the second patch information, acquired by the patch information acquisition unit 104 into multiple patch feature quantities, each representing an image feature quantity. For example, the image conversion unit 105 is configured as a vision encoder that converts image information into feature quantities represented by high-dimensional numerical vectors. For example, the image conversion unit 105 converts image information into image feature quantities after reducing the resolution of the image information representing the original image. Furthermore, for example, the image conversion unit 105 converts the image information into image features after reducing the resolution of the image information representing the original image so that the number of dimensions of the image features matches the number of divisions that divide the original image into multiple patch images.

[0022] The importance determination unit 111 determines the importance of each of the multiple patch images based on the multiple patch features acquired by the image conversion unit 105. In other words, the importance determination unit 111 determines the importance of each of the multiple patch images, including the first patch image and the second patch image, based on the multiple patch features, including the first patch features and the second patch features, acquired by the image conversion unit 105. For example, the importance determination unit 111 determines the importance of each patch image when the trained model 32 generates information based on image information by determining whether each of the multiple patch images is an important patch image with relatively high importance or an unimportant patch image with relatively low importance, based on the multiple patch features acquired by the image conversion unit 105. Alternatively, for example, the importance determination unit 111 determines the importance of each patch image when the trained model 32 generates information based on image information by determining whether the value indicating the importance of each of the multiple patch images is greater than a preset threshold, based on the multiple patch features acquired by the image conversion unit 105. Details of the process by which the importance determination unit 111 determines the importance of each patch image will be described later.

[0023] The feature extraction unit 112 extracts features corresponding to specific regions of the original image from the image features of the original image, based on the judgment result of the importance determination unit 111. For example, the feature extraction unit 112 extracts features corresponding to regions of the original image that have a higher importance than other regions from the image features of the original image, based on the judgment result of the importance determination unit 111, as important features. Specifically, the feature extraction unit 112 extracts important features from the image features of the original image that include features corresponding to important patch images, based on the judgment result of the importance determination unit 111. In other words, based on the determination result by the importance determination unit 111, the feature extraction unit 112 extracts important features that include the features corresponding to the first patch image by deleting the features corresponding to the second patch image from the image features of the original image if the importance of the first patch image is higher than the importance of the second patch image, and extracts important features that include the features corresponding to the second patch image by deleting the features corresponding to the first patch image from the image features of the original image if the importance of the second patch image is higher than the importance of the first patch image. For example, based on the determination result by the importance determination unit 111, the feature extraction unit 112 extracts important features from the image features of the original image that include the features corresponding to the important patch image and do not include the features corresponding to the unimportant patch image. In Embodiment 1, the features extracted from the image features by the feature extraction unit 112 are also called extracted features. Details of the process by which the feature extraction unit 112 extracts important features will be described later.

[0024] Next, the hardware configuration of the information processing device 100 will be described with reference to Figures 2 and 3. Figure 2 is a diagram showing an example of the hardware configuration of the information processing device 100, and Figure 3 is a diagram showing an example of the hardware configuration of the information processing device 100 that is different from Figure 2. For example, as shown in Figure 2, the information processing device 100 is composed of a computer having a processor 100a, a memory 100b, and an I / O port 100c, and is configured so that the processor 100a reads and executes a program stored in the memory 100b.

[0025] Furthermore, as shown in Figure 3, for example, the information processing device 100 is composed of a computer having a processing circuit 100d, which is dedicated hardware, and an I / O port 100c. The processing circuit 100d is composed of, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. Each function of the information processing device 100 is realized by these processors 100a or the processing circuit 100d, which is dedicated hardware, executing a program. Note that the information processing device 100 may have hardware other than that described above. Also, the hardware configuration of the model processing device 30 is the same as that of the information processing device 100, so its description is omitted.

[0026] Next, with reference to Figures 1, 4 through 7, the details of the processing performed by the information processing device 100 will be described. Figure 4 is a flowchart showing an example of the processing performed by the information processing device 100 according to Embodiment 1. The processing performed by the information processing device 100 shown in Figure 4 is the process of causing the trained model 32 to generate information in response to a request from the user regarding the original image.

[0027] As shown in Figure 4, when the information processing device 100 starts processing, it first acquires image information (step ST01). For example, in this process, the information processing device 100 acquires image information indicating the original image from the input device 10 using the image information acquisition unit 103.

[0028] Figure 5 is a diagram showing an example of a source image I1 related to image information acquired by the information processing device 100 according to Embodiment 1. As shown in Figure 5, for example, in the processing of step ST01, the information processing device 100 acquires image information showing a source image I1 including workers working in the factory.

[0029] When the information processing device 100 performs the processing in step ST01, it acquires multiple patch information (step ST02). In this process, the information processing device 100 uses the patch information acquisition unit 104 to acquire multiple patch information, which indicates each patch image when the original image I1 corresponding to the image information is divided into multiple patch images, based on the image information acquired in the processing in step ST01.

[0030] Figure 6 is a diagram showing an example of a state in which the original image I1 related to image information acquired by the information processing device 100 according to Embodiment 1 is divided. As shown in Figure 6, for example, in the process of step ST02, the information processing device 100 acquires patch information such that the number of dimensions n of the image features when the image information is converted into image features by the image conversion unit 105 corresponds to the number of patch information acquired by the patch information acquisition unit 104. Specifically, in the process of step ST02, the information processing device 100 acquires n patch information when the original image I1 is divided into n patch images such that the number of dimensions n of the image features when the image information is converted into image features by the image conversion unit 105 matches the number of patch information acquired by the patch information acquisition unit 104.

[0031] When the information processing device 100 performs the processing in step ST03, it converts the image information and each patch information into feature quantities (step ST13). In this process, the information processing device 100 converts the image information obtained in step ST01 and the multiple patch information obtained in step ST02 into image feature quantities and multiple patch feature quantities, respectively, thereby obtaining feature quantities that represent the original image I1 and the multiple patch images. For example, in this process, the information processing device 100 obtains the following feature quantities as the image feature quantity of the image information representing the original image I1 and the patch feature quantity of the patch information representing each patch image. Note that in this process, the information processing device 100 converts the image information into image feature quantities such that the number of dimensions of the image feature quantities becomes n. In other words, in this process, the information processing device 100 converts the image information into image feature quantities such that each component of the image feature quantity corresponds to each of the n patch images, as follows. Each component of the image feature quantity corresponds to the first patch feature quantity, the second patch feature quantity, ..., the nth feature quantity. Image features: [0.52, 0.07, ..., 0.94] First patch features: [0.87, 0.66, ..., 0.77] Second patch features: [0.01, 0.56, ..., 0.56] ...nth patch features: [0.12, 0.34, ..., 0.56]

[0032] When the information processing device 100 performs the processing in step ST13, it determines the importance of each patch image (step ST15). In this process, the information processing device 100 determines the importance of each of the multiple patch images corresponding to the multiple patch features obtained in the processing in step ST13. For example, in this process, the information processing device 100 calculates the similarity between each patch feature and the features of a patch image that has been set in advance to have relatively low importance, and determines the importance based on whether the calculated similarity is higher than a pre-set threshold. Also, for example, in this process, the information processing device 100 determines the importance of the multiple patch images obtained by dividing the original image into multiple patch images, based on whether the similarity between them is higher than a pre-set threshold. Patch images with high similarity to other patch images can be considered to have relatively low importance. For example, the information processing device 100 calculates the similarity between multiple features based on the Euclidean distance between the multiple features. Furthermore, for example, in this process, the information processing device 100 determines whether each patch feature is a feature that indicates that the patch image is uniform in that one or more of the brightness, saturation, and hue are uniform, and determines the importance of each patch image by considering the uniform patch image to be a patch image with relatively low importance.

[0033] When the information processing device 100 performs the processing in step ST15, it extracts important features from the image features (step ST16). In this process, based on the determination result in step ST15, the information processing device 100 extracts features corresponding to specific regions of the original image from the image features of the original image as important features using the feature extraction unit 112. For example, in this process, based on the determination result in step ST15, the information processing device 100 extracts important features from the image features of the original image that correspond to regions of the original image that have a higher importance than other regions. For example, in this process, the information processing device 100 extracts important features from the image features that include features corresponding to important patch images with a relatively high importance by deleting from the image features features that correspond to patch images that have a relatively low importance in the processing in step ST15. In other words, in this process, the information processing device 100 extracts important features from the image features, including features corresponding to patch images with relatively high importance, by replacing the values ​​of the components of the image features that correspond to patch images that were determined to have relatively low importance in the process of step ST15 with 0.

[0034] Figure 7 shows an example of a state in which a region A1 of the original image I1 with high importance has been extracted from the image information acquired by the information processing device 100 according to Embodiment 1. As shown in Figure 7, for example, if the information processing device 100 determines in step ST15 that the importance of the patch image of region A2 (hatched portion) is relatively low, in step ST16, it extracts important features, which are feature quantities corresponding to the patch image of region A1 obtained by removing region A2 from the original image I1, from the feature quantities of the original image I1.

[0035] When the information processing device 100 performs the processing in step ST16, it outputs the important features to the information conversion unit 31 of the model processing device 30 (step ST17). In this process, the information processing device 100 outputs the important features, which are the features of the high-importance region A1 of the original image I1, from the output unit 102 to the information conversion unit 31 of the model processing device 30. In order to have the trained model 32 generate information based on the important features, the information processing device 100 converts the important features into information that the trained model 32 can process.

[0036] When the model processing device 30 receives important features from the information processing device 100, the information conversion unit 31 converts the important features into information that can be processed by the trained model 32. After the information conversion unit 31 converts the important features, the model processing device 30 inputs the information generated by the conversion into the trained model 32, and the trained model 32 generates new information, which is generated information. For example, when the model processing device 30 receives important features from the information processing device 100, and also receives request information and image information from the input device 10, indicating a user's request regarding the original image in natural language, the important features and image information are input to the information conversion unit 31, and the information conversion unit 31 converts the important features and image information into information that can be processed by the trained model 32. The model processing device 30 then inputs this request information, as well as the information generated by the conversion of important features and image information by the information conversion unit 31, into the trained model 32, and the trained model 32 generates generated information in response to the user's request regarding the original image. As a result, the trained model 32 can generate generated information in response to user requests based on image features, important features, and request information, by focusing on the regions of the original image that have high importance. The model processing device 30 outputs the generated information to the display device 20, and displays the generated information on the display device 20.

[0037] The information processing device 100 terminates processing after performing the processing in step ST17. For example, the information processing device 100 is configured to perform the processing from step ST01 to step ST17 shown in Figure 4 each time image information is acquired by the image information acquisition unit 103.

[0038] As described above, the information processing device 100 according to Embodiment 1 includes: an image information acquisition unit 103 that acquires image information representing an image; a patch information acquisition unit 104 that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit 105 that converts the image information, the first patch information and the second patch information into image features, the first patch features and the second patch features, respectively; an importance determination unit 111 that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; and a feature extraction unit 112 that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit 111.

[0039] For example, the information processing device 100 according to Embodiment 1 includes an output unit 102 that outputs the extracted features extracted by the feature extraction unit 112 to an information conversion unit 31 that converts the extracted features extracted by the feature extraction unit 112 into information that can be processed by a trained model 32 that accepts input of information expressed in natural language.

[0040] Furthermore, for example, the information processing device 100 according to Embodiment 1 is configured to extract important features of high-importance regions of an image from the image features by deleting features corresponding to one of the first patch image and the second patch image included in the image features based on the determination result by the importance determination unit 111.

[0041] With this configuration, the information processing device 100 extracts features corresponding to specific regions of the original image from the image features of the original image based on the importance of each patch image when the original image is divided into multiple patch images. Therefore, even if the original image contains multiple pieces of information such as objects, people, characters, and backgrounds, when the trained model generates information based on the extracted features, it becomes possible to have the trained model focus on processing the regions with high importance, thereby suppressing hallucination when processing image information with the trained model compared to conventional methods.

[0042] In Embodiment 1, the feature extraction unit 112 is configured to extract important features from the image features of the original image that include features corresponding to important patch images and do not include features corresponding to non-important patch images, based on the determination result by the importance determination unit 111, but is not limited to this. The feature extraction unit only needs to be configured to extract important features from the image features of the original image that include features corresponding to important patch images, based on the determination result by the importance determination unit. For example, the feature extraction unit may be configured to extract features that include a portion of non-important patch images as important features from the image features of the original image, or it may be configured to extract features that do not include some of the important patch images among a plurality of important patch images as important features.

[0043] Furthermore, in Embodiment 1, the information processing system is configured to display the generated information generated by the trained model 32 using the display device 20, but it is not limited to this. The information processing system only needs to be configured to output the generated information generated by the trained model 32 as information that can be recognized by the user. For example, the information processing system may be equipped with a speaker, which is an audible signaling device, in place of or in addition to the display device, and configured to output the generated information as sound.

[0044] Furthermore, in Embodiment 1, the information processing system has the model processing device 30 including the information conversion unit 31 and the trained model 32, but the present invention is not limited thereto. The information processing device may include some or all of the functions of the model processing device. For example, when the information processing device includes the information conversion unit and the trained model, each configuration of the information processing device may itself have a function as an output unit that outputs information from each configuration of the information processing device to the information conversion unit and the trained model.

[0045] Embodiment 2. Next, an information processing system 1A according to Embodiment 2 will be described with reference to FIGS. 8 to 11. The information processing system 1A according to Embodiment 2 differs from the information processing system 1 according to Embodiment 1 in a configuration related to processing of extracting important feature amounts from a region having high image importance, but other configurations are the same. Configurations similar to those in Embodiment 1 are given the same names and reference signs as in Embodiment 1, and descriptions thereof are omitted.

[0046] FIG. 8 is a block diagram showing a schematic configuration of the information processing system 1A according to Embodiment 2. As shown in FIG. 8, the information processing system 1A according to Embodiment 2 includes an input device 10, a display device 20, a model processing device 30, and an information processing device 100A, which are connected wirelessly or by wire so that they can communicate with each other. Note that the input device 10, the display device 20, the model processing device 30, and the information processing device 100A may be communicably connected to each other via another device (not shown) or a communication line.

[0047] The information processing device 100A includes an input unit 101, an output unit 102, an image information acquisition unit 103, a patch information acquisition unit 104, an image conversion unit 105, a request acquisition unit 106, a noun information extraction unit 108, a character string conversion unit 109, a similarity calculation unit 110, an importance determination unit 111A, a feature amount extraction unit 112, and a generated information acquisition unit 113.

[0048] The request acquisition unit 106 acquires request information indicating a request related to an original image from a user in natural language. For example, the request acquisition unit 106 acquires request information that indicates a request related to an original image from a user in natural language, such as an answer to a question about the original image from the user or a request for provision of information about the original image from the user. For example, the request acquisition unit 106 acquires request information from the input device 10 via the input unit 101

[0049] The noun information extraction unit 108 extracts noun information, which is character string information indicating nouns included in the request information, by performing morphological analysis on the request information acquired by the request acquisition unit 106. The noun information extraction unit 108 may be configured to extract noun information corresponding to each noun for all nouns included in the request information by morphological analysis of the request information acquired by the request acquisition unit 106. For a specific noun included in the request information, for example, when the request information includes a plurality of nouns, the noun information extraction unit 108 may be configured to extract noun information corresponding to any one or more of the plurality of nouns, may be configured to extract noun information corresponding to any one or more subjects among the plurality of nouns, may be configured to extract noun information corresponding to any one or more objects among the plurality of nouns, or may be configured to extract corresponding noun information for each of one or more subjects and objects.

[0050] The character string conversion unit 109 acquires a noun feature quantity indicating a noun included in the request information by converting the noun information extracted by the noun information extraction unit 108 into a noun feature quantity. For example, the character string conversion unit 109 is configured by a Text Encoder that converts information indicated by a character string including a noun into a feature quantity indicated by a high-dimensional numerical vector.

[0051] The similarity calculation unit 110 calculates the similarity between the noun feature obtained by the conversion in the string conversion unit 109 and each patch feature. In other words, the similarity calculation unit 110 calculates the similarity between the noun feature obtained by the string conversion unit 109 and each patch feature obtained by the image conversion unit 105, including a first similarity between the noun feature obtained by the string conversion unit 109 and the first patch feature obtained by the image conversion unit 105, and a second similarity between the noun feature and the second patch feature obtained by the image conversion unit 105. For example, the similarity calculation unit 110 calculates the Euclidean distance as the similarity between the numerical vector representing the noun feature and the numerical vector representing each patch feature. The similarity calculation unit 110 may be configured to calculate cosine similarity, Hamming distance, Manhattan distance, or Pearson correlation coefficient instead of Euclidean distance as the similarity between the numerical vector representing the noun feature and the numerical vector representing each patch feature. Furthermore, the similarity calculation unit may consist of a large-scale language model that calculates the similarity between the input noun features and each patch feature based on the input noun features and multiple patch features.

[0052] Furthermore, if there are multiple noun pieces of information extracted by the noun information extraction unit 108, the similarity calculation unit 110 may be configured to calculate the similarity for all combinations of each noun piece of information and each patch feature, or it may be configured to calculate the similarity only between any one of the noun pieces of information and each patch feature.

[0053] The importance determination unit 111A determines the importance of each of the multiple patch images acquired by the patch information acquisition unit 104 based on the similarity between the noun feature calculated by the similarity calculation unit 110 and each patch feature. In other words, the importance determination unit 111A determines the importance of each of the multiple patch images, including the first patch image and the second patch image, based on the similarity between the noun feature and each patch feature, including the first and second similarity calculated by the similarity calculation unit 110.

[0054] For example, the importance determination unit 111A determines whether the patch image corresponding to a specific patch feature is an important patch image with relatively high importance or an unimportant patch image with relatively low importance, based on the similarity between the noun feature calculated by the similarity calculation unit 110 and the specific patch feature. Specifically, the importance determination unit 111A determines whether the patch image corresponding to the specific patch feature is an important patch image with relatively high importance or an unimportant patch image with relatively low importance, based on whether the similarity between the noun feature calculated by the similarity calculation unit 110 and the specific patch feature is higher than a preset threshold. In other words, the importance determination unit 111A determines that if the similarity between the noun feature calculated by the similarity calculation unit 110 and the specific patch feature is higher than a preset threshold, the patch image corresponding to the specific patch feature is an important patch image with relatively high importance. Conversely, if the similarity between the noun feature calculated by the similarity calculation unit 110 and the specific patch feature is less than or equal to a preset threshold, the importance determination unit 111A determines that the patch image corresponding to the specific patch feature is an unimportant patch image with relatively low importance. As described above, since the similarity calculation unit 110 calculates the first and second similarities based on the noun feature, the first patch feature, and the second patch feature, it can be said that the importance determination unit 111A determines the importance of the first patch image and the second patch image based on the first and second patch features.

[0055] The generated information acquisition unit 113 acquires generated information generated by the trained model 32 based on the input of important features converted by the information conversion unit 31 from the model processing device 30 via the input unit 101.

[0056] The hardware configuration of the information processing device 100A is the same as that of the information processing device 100 according to Embodiment 1, so its description will be omitted.

[0057] Next, with reference to Figures 8 to 11, the details of the processing performed by the information processing device 100A will be described. Figure 9 is a flowchart showing an example of the processing performed by the information processing device 100A according to Embodiment 2. Note that some of the processing performed by the information processing device 100A according to Embodiment 2 is the same as the processing performed by the information processing device 100 according to Embodiment 1, so the same processing as in Embodiment 1 is denoted by the same reference numerals as in Embodiment 1 and its description is omitted.

[0058] As shown in Figure 9, when the information processing device 100A starts processing, it first acquires image information (step ST01). After performing the processing in step ST01, the information processing device 100A acquires multiple patch information (step ST02).

[0059] When the information processing device 100A performs the processing in step ST02, it acquires request information (step ST03). In this process, the information processing device 100A acquires request information from the input device 10 that indicates the user's request regarding the original image in natural language. For example, if the original image is an image that includes workers working in a factory, the information processing device 100A acquires request information that is indicated by the string "What are the workers doing?".

[0060] When the information processing device 100A performs the processing in step ST03, it extracts noun information contained in the request information (step ST11). In this process, the information processing device 100A extracts noun information indicating the nouns contained in the request information by performing morphological analysis on the request response obtained in the processing of step ST03.

[0061] Figure 10 shows an example of the results of morphological analysis by the information processing device 100A according to Embodiment 2. As shown in Figure 10, for example, when the information processing device 100A obtains request information represented by the string "What is the worker doing?" in step ST03, in step ST11, it subdivides the string into morphemes and extracts the noun "worker" from each morpheme.

[0062] When the information processing device 100A performs the processing in step ST11, it converts the noun information into noun features (step ST12). In this process, the information processing device 100A converts the noun information extracted in step ST11 into noun features using the string conversion unit 109. For example, in this process, the information processing device 100A converts the noun information extracted in step ST11, such as "worker," into a noun feature represented by a high-dimensional numerical vector.

[0063] When the information processing device 100A performs the processing in step ST12, it converts the image information and each patch information into feature quantities (step ST13).

[0064] When the information processing device 100A performs the processing in step ST13, it calculates the similarity between the noun feature and each patch feature (step ST14). In this process, the information processing device 100A calculates the similarity between the noun feature obtained in step ST12 and each patch feature obtained in step ST13 using the similarity calculation unit 110.

[0065] When the information processing device 100A performs the processing in step ST14, it determines the importance of each patch image (step ST25). In this process, the information processing device 100A determines the importance of each of the multiple patch images corresponding to the multiple patch information obtained in the processing of step ST02, based on the similarity calculated in the processing of step ST14. For example, in this process, the information processing device 100A determines the importance of each patch image based on whether the similarity between the calculated noun feature and each patch feature is higher than a preset threshold.

[0066] When the information processing device 100A performs the processing in step ST25, it extracts important features from the image features (step ST16).

[0067] When the information processing device 100A performs the processing in step ST16, it outputs the image features and important features to the information conversion unit 31 of the model processing device 30 (step ST27). In this process, the information processing device 100A outputs the image features of the original image and the important features, which are the features of the regions of the original image that are of high importance, from the output unit 102 to the information conversion unit 31, thereby converting the image features and important features into information that the trained model 32 can process.

[0068] When the information processing device 100A performs the processing in step ST27, it outputs the request information to the trained model 32 (step ST18). In this process, the information processing device 100A outputs the request information to the trained model 32 of the model processing device 30 via the output unit 102 in order to cause the trained model 32 to generate information based on the image features, important features, and request information. When the model processing device 30 receives the image features, important features, and request information from the information processing device 100A, the information conversion unit 31 converts the image features and important features into information that the trained model 32 can process. After the information conversion unit 31 converts the image features and important features, the model processing device 30 inputs the information generated by the conversion of the image features and important features, as well as the request information, to the trained model 32, and the trained model 32 generates generated information in response to the user's request regarding the original image. As a result, the trained model 32 can generate generated information in response to user requests based on image features, important features, and request information, by focusing on the regions of the original image that have high importance when generating generated information.

[0069] When the information processing device 100A performs the processing in step ST18, it acquires generated information from the trained model 32 (step ST19). In this process, the information processing device 100A acquires generated information that was generated by the trained model 32 based on the information input to the model processing device 30 in steps ST27 and ST18.

[0070] When the information processing device 100A performs the processing in step ST19, it outputs the generated information to the display device 20 (step ST20). In this process, the information processing device 100A outputs the generated information acquired in the processing of step ST19 to the display device 20 via the output unit 102 in order to display it on the display device 20. The display device 20 displays the input generated information.

[0071] The information processing device 100A terminates processing after performing the processing in step ST20. For example, the information processing device 100A is configured to perform the processing from step ST01 to step ST20 shown in Figure 9 each time image information is acquired by the image information acquisition unit 103.

[0072] As described above, the information processing device 100A according to Embodiment 2 includes an image information acquisition unit 103 that acquires image information representing an image, a patch information acquisition unit 104 that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information, an image conversion unit 105 that converts the image information, the first patch information and the second patch information into image features, the first patch features and the second patch features, respectively, an importance determination unit 111A that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features, and an importance determination unit The system includes a feature extraction unit 112 that extracts features corresponding to a specific region of an image from the image features based on the determination result by 111A, an output unit 102 that outputs the features extracted by the feature extraction unit 112 to an information conversion unit 31 that converts the extracted features extracted by the feature extraction unit 112 into information that can be processed by a trained model 32 that accepts input of information expressed in natural language, and a generated information acquisition unit 113 that acquires generated information generated by the trained model 32 based on the input of the extracted features converted by the information conversion unit 31, and is configured to output the generated information to the display device 20 so that the generated information is displayed on the display device 20.

[0073] With this configuration, the information processing device 100A can display on the display device 20 the information generated by the trained model 32 based on the features extracted from the image features by the feature extraction unit 112.

[0074] Furthermore, the information processing device 100A according to Embodiment 2 includes a request acquisition unit 106 that acquires request information indicating a user's request for an image in natural language, and is configured to determine the importance of the first patch image and the second patch image based on the first patch feature quantity, the second patch feature quantity, and the request information.

[0075] For example, the information processing device 100A according to Embodiment 2 includes a noun information extraction unit 108 that extracts noun information indicating nouns included in the request information by morphological analysis of the request information, and a string conversion unit 109 that converts the noun information into noun features, and is configured to determine the importance of the first patch image and the second patch image based on the first patch features, the second patch features, and the noun features.

[0076] With this configuration, the information processing device 100A can have the trained model 32 focus on processing nouns, which are relatively more important than parts of speech such as particles and conjunctions, among the information included in the requested information. This makes it possible to suppress hallucination when the trained model 32 processes image information more effectively than before.

[0077] Furthermore, for example, the information processing device 100A according to Embodiment 2 includes a similarity calculation unit 110 that calculates a first similarity between a noun feature and a first patch feature, and a second similarity between a noun feature and a second patch feature, and is configured to determine the importance of the first patch image and the second patch image based on the first and second similarities.

[0078] With this configuration, the information processing device 100A extracts features corresponding to a specific region of the original image from the image features of the original image based on the similarity between the noun information included in the request information and the patch information representing each patch image when the original image is divided into multiple patch images. Therefore, when the trained model generates information based on the extracted features, it becomes possible to have the trained model focus on processing regions of high importance, thereby suppressing hallucination when processing image information with the trained model compared to conventional methods.

[0079] In Embodiment 2, the information processing system 1A is configured such that the information processing device 100A acquires generated information generated by the trained model 32 from the model processing device 30 and outputs the acquired generated information to the display device 20, but is not limited to this. The information processing system only needs to be configured to display the generated information generated by the trained model 32 on the display device 20. For example, the information processing system may be configured so that the generated information generated by the trained model is output from the model processing device to the display device 20 without going through the information processing device, thereby allowing the generated information to be displayed on the display device 20.

[0080] Embodiment 3. Next, the information processing system 1B according to Embodiment 3 will be described with reference to Figures 12 to 14. The information processing system 1B according to Embodiment 3 differs from the information processing system 1A according to Embodiment 2 in the configuration of the model processing device and some of the information processing devices. However, the same configurations as in Embodiment 2 will be given the same names and reference numerals as in Embodiment 2 and will not be described.

[0081] Figure 12 is a block diagram showing the schematic configuration of the information processing system 1B according to Embodiment 3. As shown in Figure 12, the information processing system 1B according to Embodiment 3 comprises an input device 10, a display device 20, a model processing device 30B, and an information processing device 100B, which are connected wirelessly or by wire to enable communication between them. The input device 10, the display device 20, the model processing device 30B, and the information processing device 100B may also be connected to each other via other devices or communication lines (not shown) to enable information communication between them.

[0082] The model processing device 30B includes an information conversion unit 31, a first trained model 32B, and a second trained model 33B. The information conversion unit 31 converts the information input to the model processing device 30 into information that can be processed by the first trained model 32B. The details of the information conversion unit 31 and the first trained model 32B are the same as those of the information conversion unit 31 and trained model 32 in Embodiment 1, so their explanation is omitted.

[0083] The second pre-trained model 33B accepts input string information and image information expressed in natural language and generates information corresponding to the input information. For example, the second pre-trained model 33B accepts input string information expressed in natural language and image information and calculates the similarity between the input string information and the image information. For example, the second pre-trained model 33B is composed of an open vocabulary object detection model that accepts input string information and image information and calculates the similarity between the features of the input string information and the features of the image information.

[0084] The information processing device 100B includes an input unit 101, an output unit 102, an image information acquisition unit 103, a patch information acquisition unit 104, an image conversion unit 105, a request acquisition unit 106, an instruction information generation unit 107, a noun information extraction unit 108, a string conversion unit 109, a similarity calculation unit 110, an importance determination unit 111B, a feature extraction unit 112, and a generated information acquisition unit 113.

[0085] The patch information acquisition unit 104 acquires multiple patch information, each representing a patch image when an image corresponding to the image information is divided into multiple patch images, based on the image information acquired by the image information acquisition unit 103. The function of the patch information acquisition unit 104 to acquire multiple patch information is the same as that of the patch information acquisition unit 104 in Embodiment 1, so a description is omitted. The patch information acquisition unit 104 outputs the acquired multiple patch information to the information conversion unit 31 of the model processing device 30B and the second trained model 33B via the output unit 102.

[0086] The instruction information generation unit 107 generates instruction information that instructs the first trained model 32B to generate noun information indicating nouns included in the request information, based on the request information acquired by the request acquisition unit 106. For example, the instruction information generation unit 107 generates instruction information that instructs the first trained model 32B to generate at least one of subject information, which is noun information indicating the subject included in the request information, and object information, which is noun information indicating the object included in the request information, based on the request information acquired by the request acquisition unit 106.

[0087] In Embodiment 3, the instruction information indicating instructions for the first trained model 32B to generate subject information indicating the subject included in the request information is also referred to as a "subject generation prompt." The subject generation prompt only needs to include instructions for the first trained model 32B to generate subject information, and may further include other instructions. For example, the subject generation prompt may include instructions for the first trained model 32B to generate subject information indicating all subjects included in the request information, or it may include instructions for the first trained model 32B to generate subject information indicating the subject that appears most frequently in the request information. Furthermore, the subject generation prompt may include instructions for the first trained model 32B to generate subject information indicating the first subject that appears in the sentence constituting the request information, or it may include instructions for the first trained model 32B to generate subject information indicating the last subject that appears in the sentence constituting the request information. Furthermore, the subject generation prompt may include an instruction to cause the first trained model 32B to generate subject information indicating the subject included in the request information, along with object information indicating the object included in the request information, or it may include an instruction to cause the first trained model 32B to generate other information along with subject information indicating the subject included in the request information. The instruction information generation unit 107 outputs the generated instruction information to the first trained model 32B of the model processing device 30B via the output unit 102.

[0088] When instruction information is input from the information processing device 100B, the model processing device 30B inputs the instruction information into the first trained model 32B, causing the first trained model 32B to generate noun information corresponding to the instruction information. The model processing device 30B also inputs the noun information generated by the first trained model 32B and multiple patch information from the patch information acquisition unit 104 into the second trained model 33B, causing the second trained model 33B to calculate the similarity between the input noun information and each patch information. In other words, the model processing device 30B inputs the noun information generated by the first trained model 32B and multiple patch information including the first and second patch information from the patch information acquisition unit 104 into the second trained model 33B, causing the second trained model 33B to calculate the similarity between the input noun information and each patch information, including the third similarity between the input noun information and the first patch information, and the fourth similarity between the input noun information and the second patch information.

[0089] Specifically, the model processing unit 30B inputs noun information generated by the first trained model 32B, and a plurality of patch information including the first patch information and the second patch information from the patch information acquisition unit 104, into the second trained model 33B, and causes the second trained model 33B to calculate the similarity between the features of the input noun information and the features of each patch information, including a third similarity between the features of the input noun information and the features of the first patch information, and a fourth similarity between the features of the input noun information and the features of the second patch information.

[0090] For example, the model processing unit 30B instructs the second trained model 33B to calculate the similarity between the feature quantities of noun information and the feature quantities of each patch information for each patch information, with a value between 0 and 1. Specifically, the model processing unit 30B instructs the second trained model 33B to calculate the similarity between the feature quantities of subject information generated by the first trained model 32B and the feature quantities of each patch information for each patch information, with a value between 0 and 1. The model processing unit 30B outputs the similarity between the noun information and the patch information calculated by the second trained model 33B to the information processing unit 100B.

[0091] The importance determination unit 111B determines the importance of each of the multiple patch images based on the similarity between the noun feature calculated by the similarity calculation unit 110 and each patch feature. The function of the importance determination unit 111B in determining the importance of each of the multiple patch images based on the similarity between the noun feature calculated by the similarity calculation unit 110 and each patch feature is the same as the function of the importance determination unit 111A in Embodiment 2, so a description is omitted.

[0092] Furthermore, the importance determination unit 111B determines the importance of each of the multiple patch images based on the similarity between the noun information calculated by the second trained model 33B and each patch information. In other words, the importance determination unit 111B determines the importance of each of the multiple patch images based on the similarity between the noun information and the multiple patch information, including the first similarity between the noun information and the first patch information, and the second similarity between the noun information and the second patch information, calculated by the second trained model 33B.

[0093] For example, the importance determination unit 111B determines whether each of the multiple patch images is an important patch image with relatively high importance or an unimportant patch image with relatively low importance, based on the similarity between the noun information calculated by the second trained model 33B and each patch information. Specifically, the importance determination unit 111B determines the importance of each of the multiple patch images based on whether the similarity value between the noun information calculated by the second trained model 33B and each patch information is greater than a preset threshold. In other words, if the similarity value between the subject information calculated by the second trained model 33B and each patch information is greater than a preset threshold, the importance determination unit 111B determines that the patch image corresponding to that patch information is an important patch image with relatively high importance, and if it is less than or equal to the threshold, it determines that the patch image corresponding to that patch information is an unimportant patch image with relatively low importance.

[0094] The information extraction unit 112B, acting as a feature extraction unit, extracts features corresponding to specific regions of the original image from the image features of the original image as important features, based on the determination result by the importance determination unit 111B. The function of the information extraction unit 112B in extracting features corresponding to specific regions of the original image from the image features of the original image as important features, based on the determination result by the importance determination unit 111, is the same as that of the feature extraction unit 112 in Embodiment 1, so a detailed explanation is omitted.

[0095] Furthermore, the information extraction unit 112B extracts some or all of the patch information from the multiple patch information acquired by the patch information acquisition unit 104 based on the determination result by the importance determination unit 111B. For example, the information extraction unit 112B extracts patch information corresponding to an important patch image from the multiple patch information acquired by the patch information acquisition unit 104 based on the determination result by the importance determination unit 111B. Also, for example, the information extraction unit 112B extracts patch information corresponding to an important patch image and patch images that are adjacent to the important patch image in the original image from the multiple patch information acquired by the patch information acquisition unit 104 based on the determination result by the importance determination unit 111B.

[0096] The hardware configuration of the information processing device 100B and the model processing device 30B is the same as that of the information processing device 100 according to Embodiment 1, so a description will be omitted.

[0097] Next, with reference to Figures 12 to 14, the details of the processing performed by the information processing device 100B will be described. Figure 13 is a flowchart showing an example of the processing performed by the information processing device 100B according to Embodiment 3. Note that some of the processing performed by the information processing device 100B according to Embodiment 3 is the same as the processing performed by the information processing device 100 according to Embodiment 1, so the same processing as in Embodiment 1 is denoted by the same reference numerals as in Embodiment 1 and its description is omitted.

[0098] As shown in Figure 13, when the information processing device 100B starts processing, it first acquires image information (step ST01).

[0099] When the information processing device 100B performs the processing in step ST01, it acquires multiple patch information (step ST02). When the information processing device 100B performs the processing in step ST02, it acquires request information (step ST03).

[0100] When the information processing device 100B performs the processing in step ST03, it generates a subject generation prompt (step ST04). For example, if the information processing device 100B has obtained request information from the user in the processing of step ST03, indicated by the string "What is the worker doing?", then in the processing of step ST04, it generates a subject generation prompt, which is instruction information indicated by the following string: "Please extract the noun that will be the subject of the text from the following sentence. Output format should be nouns only. "What is the worker doing?""

[0101] When the information processing device 100B performs the processing in step ST04, it outputs a subject generation prompt to the first trained model 32B (step ST05). In this process, the information processing device 100B outputs the subject generation prompt generated in step ST04 to the first trained model 32B of the model processing device 30, thereby causing the first trained model 32B to extract the subject included in the request information.

[0102] When the model processing device 30B receives a subject generation prompt from the information processing device 100B, it inputs the input subject generation prompt to the first trained model 32B, causing the first trained model 32B to extract subject information included in the request information. For example, based on the input of the subject generation prompt, the first trained model 32B extracts subject information indicating the subject "worker" from the request information. Once the subject information has been extracted by the first trained model 32B, the model processing device 30B outputs the extracted subject information to the information processing device 100B.

[0103] When the information processing device 100B performs the processing in step ST05, it obtains subject information from the first trained model 32B (step ST06). In this process, the information processing device 100B obtains subject information generated by the first trained model 32B based on the subject generation prompt output in the processing of step ST05 from the model processing device 30 using the input unit 101.

[0104] When the information processing device 100B performs the processing in step ST06, it outputs subject information and multiple patch information to the second trained model 33B (step ST08). In this process, the information processing device 100B outputs this subject information and multiple patch information to the second trained model 33B of the model processing device 30B in order to calculate the similarity between the subject information obtained in step ST06 and the patch information obtained in step ST02. The information processing system 1B may be configured so that the subject information generated by the first trained model 32B is input to the second trained model 33B without going through the information processing device 100B. In such a case, the information processing system 1B may be configured so that the multiple patch information is output from the information processing device 100B to the second trained model 33B when the subject generation prompt is output from the information processing device 100B to the first trained model 32B in the processing of step ST05.

[0105] The model processing unit 30B inputs the subject information and multiple patch information received from the information processing unit 100B in step ST08 into the second trained model 33B, causing the second trained model 33B to calculate the similarity between the subject information and each patch information. The model processing unit 30B outputs the similarity between the subject information and each patch information calculated by the second trained model 33B to the information processing unit 100B.

[0106] When the information processing device 100B performs the processing in step ST08, it obtains the similarity between the subject information and each patch information from the second trained model 33B (step ST09). In this process, the information processing device 100B obtains the similarity between the subject information output to the second trained model 33B in the processing of step ST08 from the second trained model 33B of the model processing device 30B.

[0107] Figure 14A is a diagram showing an example of the similarity between subject information and each patch information acquired by the information processing device 100B according to Embodiment 3. As shown in Figure 14A, for example, in the processing of step ST09, the information processing device 100B acquires the similarity between subject information and each patch information for each patch information, with a value ranging from 0 to 1. Each value shown in Figure 14A represents the similarity between subject information and each patch information, and the position of each value indicates the position of the patch image in the original image corresponding to each similarity.

[0108] When the information processing device 100B performs the processing in step ST09, it extracts some patch information based on the similarity between the subject information and each patch information (step ST10). In this process, the information processing device 100B uses the information extraction unit 112B to extract some patch information from the multiple patch information obtained in the processing of step ST02, based on the similarity between the subject information obtained in the processing of step ST09 and each patch information.

[0109] Figure 14B is a diagram showing an example of a state in which some patch information has been extracted from a plurality of patch information acquired by the information processing device 100B according to Embodiment 3 based on the similarity between the subject information and the patch information. For example, in the process of step ST10, the information processing device 100B uses the information extraction unit 112B to extract, from the plurality of patch information acquired in the process of step ST02, patch information whose similarity to the subject information is greater than a preset similarity threshold of 0.6, and patch information corresponding to patch images that are adjacent to the patch image in the original image, i.e., patch information corresponding to the similarity enclosed by the solid shown in Figure 14B, based on the similarity between the subject information shown in Figure 14A acquired in the process of step ST09 and each patch information. In other words, as shown in Figure 14B, in the process of step ST10, the information processing device 100B uses the information extraction unit 112B to extract patch information corresponding to similarity scores of 0.7, 0.8, and 0.9 between the subject information and each patch information, and patch information corresponding to patch images adjacent to the patch images corresponding to these patch information in the original image.

[0110] When the information processing device 100B performs the processing in step ST10, it extracts noun information contained in the request information (step ST11). In this process, the information processing device 100B may extract new noun information contained in the request information, or it may be configured to perform subsequent processing using the subject information obtained in the processing in step ST06 without extracting new noun information. For example, when the information processing device 100B extracts new noun information contained in the request information, it may be configured to extract noun information that indicates nouns other than the subject contained in the request information.

[0111] When the information processing device 100B performs the processing in step ST11, it converts noun information into noun features (step ST12).

[0112] When the information processing device 100B performs the processing in step ST12, it converts the image information and each patch information into feature quantities (step ST23). In this process, the information processing device 100B converts the image information acquired in step ST01 and the each patch information extracted in step ST10 into feature quantities, respectively.

[0113] When the information processing device 100B performs the processing in step ST23, it calculates the similarity between the noun feature and each patch feature (step ST14). When the information processing device 100B performs the processing in step ST14, it determines the importance of each patch image (step ST25). When the information processing device 100B performs the processing in step ST25, it extracts important features from the image features (step ST16). When the information processing device 100B performs the processing in step ST16, it outputs the image features and important features to the information conversion unit 31 of the model processing device 30B (step ST27). In this process, the information processing device 100B outputs the image features of the original image and the important features, which are the features of the regions with high importance in the original image, from the output unit 102 to the information conversion unit 31, thereby converting the image features and important features into information that the first trained model 32B can process.

[0114] When the information processing device 100B performs the processing in step ST27, it outputs the request information to the first trained model 32B (step ST28). In this process, the information processing device 100B outputs the request information to the first trained model 32B of the model processing device 30B via the output unit 102 in order to cause the first trained model 32B to generate information based on the image features, important features, and request information. When the model processing device 30B receives the image features, important features, and request information from the information processing device 100B, the information conversion unit 31 converts the image features and important features into information that the first trained model 32B can process. When the information conversion unit 31 converts the image features and important features, the model processing device 30B inputs the information generated by the conversion of the image features and important features, as well as the request information, to the first trained model 32B, and the first trained model 32B generates generated information in response to the user's request regarding the original image. As a result, the first trained model 32B can generate generated information in response to user requests based on image features, important features, and request information, by focusing on the regions of the original image that have high importance.

[0115] When the information processing device 100B performs the processing in step ST28, it acquires generated information from the first trained model 32B (step ST29). In this process, the information processing device 100B acquires generated information that was generated by the first trained model 32B based on the information input to the model processing device 30B in steps ST27 and ST18.

[0116] When the information processing device 100B performs the processing in step ST29, it outputs the generated information to the display device 20 (step ST20). In this process, the information processing device 100B outputs the generated information acquired in the processing of step ST19 to the display device 20 via the output unit 102 in order to display it on the display device 20. The display device 20 displays the input generated information.

[0117] The information processing device 100B terminates processing after performing the processing in step ST20. For example, the information processing device 100B is configured to perform the processing from step ST01 to step ST20 shown in Figure 13 each time image information is acquired by the image information acquisition unit 103.

[0118] As described above, the information processing device 100B according to Embodiment 3 comprises: an image information acquisition unit 103 that acquires image information representing an image; a patch information acquisition unit 104 that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit 105 that converts the image information, the first patch information and the second patch information into image features, first patch features and second patch features, respectively; an importance determination unit 111B that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; a feature extraction unit 112 that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit 111B; and a request acquisition unit 106 that acquires request information that indicates a user's request regarding an image in natural language. The device is configured to determine the importance of the first patch image and the second patch image based on the first patch features, the second patch features and the request information.

[0119] For example, the information processing device 100B according to Embodiment 3 includes an instruction information generation unit 107 that generates instruction information indicating an instruction for the first trained model 32B to generate noun information indicating nouns included in the request information based on the request information, and a string conversion unit 109 that converts the noun information generated by the first trained model 32B into noun features, and is configured to determine the importance of the first patch image and the second patch image based on the first patch features, the second patch features and the noun features.

[0120] With this configuration, the information processing device 100B can have the trained model focus on processing nouns, which are relatively more important than parts of speech such as particles and conjunctions, among the information included in the requested information. This makes it possible to suppress hallucination when processing image information with the trained model compared to conventional methods.

[0121] The information processing system 1B according to Embodiment 3 includes, but is not limited to, an information conversion unit 31, a first learned model 32B, and a second learned model 33B, which convert information input to the model processing device 30 into information that can be processed by the first learned model 32B. For example, the information processing system may include one learned model having the functions of the information conversion unit, the first learned model, and the second learned model; or it may include a first learned model and one learned model having the functions of the information conversion unit and the second learned model; or any other configuration may have some of the functions of the information conversion unit, the first learned model, and the second learned model, and the functions of the model processing device may be realized through the cooperation of these information conversion unit, the first learned model, and the second learned model; or each of a plurality of independently formed devices may be configured to include any of the information conversion unit, the first learned model, and the second learned model.

[0122] Furthermore, the first trained model and the second trained model may each be composed of multiple trained models. For example, the trained model that generates a subject generation prompt in step ST04, the trained model that generates subject information based on the input of the subject generation prompt generated in step ST04, and the trained model that extracts noun information in step ST11 may each be multiple trained models that function independently of each other. These multiple trained models may be possessed by multiple devices formed independently of each other, or the function of one of these trained models may be realized through the cooperation of multiple devices.

[0123] Furthermore, the information processing device 100B according to Embodiment 3 is configured such that the instruction information generation unit 107 generates a subject generation prompt, which is an instruction to cause the first trained model 32B to generate subject information indicating the subject included in the request information. However, the device is not limited to this configuration. For example, the information processing device may be configured such that, instead of a subject generation prompt, the instruction information generation unit generates an object generation prompt, which is an instruction to cause the first trained model to generate object information indicating the object included in the request information.

[0124] Embodiment 4. Next, the information processing system 1C according to Embodiment 4 will be described with reference to Figures 12 and 15. The information processing system 1C according to Embodiment 4 differs from the information processing system 1B according to Embodiment 3 in the configuration of the model processing device and some of the information processing devices. However, the same configurations as in Embodiment 3 will be given the same names and reference numerals as in Embodiment 3 and will not be described.

[0125] Figure 12 is a block diagram showing the schematic configuration of the information processing system 1C according to Embodiment 4. As shown in Figure 12, the information processing system 1C according to Embodiment 4 comprises an input device 10, a display device 20, a model processing device 30C, and an information processing device 100C, which are connected wirelessly or by wire to enable communication between them. The input device 10, the display device 20, the model processing device 30C, and the information processing device 100C may also be connected to each other via other devices or communication lines (not shown) to enable information communication between them.

[0126] The model processing device 30C includes an information conversion unit 31, a first trained model 32B, and a second trained model 33C.

[0127] The second pre-trained model 33C accepts input string information and image information expressed in natural language and generates information corresponding to the input information. For example, the second pre-trained model 33C accepts input string information expressed in natural language strings and image information and calculates the similarity between the input string information and the image information. For example, the second pre-trained model 33C is composed of a VLM (Vision and Language Model) that accepts input string information and image information and calculates the similarity between the input string information and the image information.

[0128] The information processing device 100C includes an input unit 101, an output unit 102, an image information acquisition unit 103, a patch information acquisition unit 104, an image conversion unit 105, a request acquisition unit 106, an instruction information generation unit 107C, a noun information extraction unit 108, a string conversion unit 109, a similarity calculation unit 110, an importance determination unit 111C, a feature quantity extraction unit 112, and a generated information acquisition unit 113.

[0129] The instruction information generation unit 107C generates instruction information that instructs the first trained model 32B to generate noun information indicating nouns included in the request information, based on the request information acquired by the request acquisition unit 106. For example, the instruction information generation unit 107C generates instruction information that instructs the first trained model 32B to generate at least one of subject information, which is noun information indicating the subject included in the request information, and object information, which is noun information indicating the object included in the request information, based on the request information acquired by the request acquisition unit 106.

[0130] Furthermore, the instruction information generation unit 107C generates a similarity calculation prompt, which is instruction information that indicates in natural language an instruction for the second trained model 33C to calculate the similarity between the noun information and each patch information acquired by the patch information acquisition unit 104, based on the noun information generated by the first trained model 32B. For example, the instruction information generation unit 107C generates a similarity calculation prompt, which is instruction information that indicates an instruction for the second trained model 33C to calculate the similarity between the subject information and each patch information acquired by the patch information acquisition unit 104, based on the subject information generated by the first trained model 32B. The instruction information generation unit 107C outputs the generated similarity calculation prompt to the second trained model 33C of the model processing device 30C via the output unit 102.

[0131] When the model processing unit 30C receives instruction information from the information processing unit 100C, it inputs the instruction information to the first trained model 32B and the second trained model 33C, causing the first trained model 32B and the second trained model 33C to generate information corresponding to the instruction information. For example, when the model processing unit 30B receives a subject generation prompt from the information processing unit 100C, it inputs the subject generation prompt to the first trained model 32B, causing the first trained model 32B to generate subject information corresponding to the subject generation prompt.

[0132] Furthermore, for example, when the model processing device 30C receives a similarity calculation prompt and multiple patch information acquired by the patch information acquisition unit 104 from the information processing device 100C, it inputs the similarity calculation prompt and these multiple patch information to the second trained model 33C, causing the second trained model 33C to calculate the similarity between the noun information and each patch information. In other words, when the model processing device 30C receives a similarity calculation prompt and multiple patch information, including the first patch information and the second patch information acquired by the patch information acquisition unit 104, from the information processing device 100C, it inputs the similarity calculation prompt and multiple patch information, including the first patch information and the second patch information, to the second trained model 33C, causing the second trained model 33C to calculate the similarity between the input noun information and each patch information, including the third similarity between the input noun information and the first patch information, and the fourth similarity between the input noun information and the second patch information.

[0133] Specifically, when the model processing device 30C receives a similarity calculation prompt and multiple patch information, including the first patch information and the second patch information acquired by the patch information acquisition unit 104, from the information processing device 100C, it inputs the similarity calculation prompt and the multiple patch information, including the first patch information and the second patch information, into the second trained model 33C, causing the second trained model 33C to calculate the similarity between the input noun information and each patch information, including the third similarity between the feature quantity of the input noun information and the feature quantity of the first patch information, and the fourth similarity between the feature quantity of the input noun information and the feature quantity of the second patch information.

[0134] For example, the model processing unit 30C instructs the second trained model 33C to calculate the similarity between noun information and each patch information for each patch information, with a value between 0 and 1. Specifically, the model processing unit 30C instructs the second trained model 33C to calculate the similarity between the feature quantities of subject information generated by the first trained model 32B and the feature quantities of each patch information for each patch information, with a value between 0 and 1. The model processing unit 30C outputs the similarity between noun information and patch information calculated by the second trained model 33C to the information processing unit 100C.

[0135] The importance determination unit 111C determines the importance of each of the multiple patch images based on the similarity between the noun feature calculated by the similarity calculation unit 110 and each patch feature. The function of the importance determination unit 111B, which determines the importance of each of the multiple patch images based on the similarity between the noun feature calculated by the similarity calculation unit 110 and each patch feature, is the same as the function of the importance determination unit 111A in Embodiment 2, so its explanation is omitted.

[0136] Furthermore, the importance determination unit 111C determines the importance of each of the multiple patch images based on the similarity between the noun information calculated by the second trained model 33C and each patch information. For example, the importance determination unit 111C determines the importance of each of the multiple patch images based on whether the similarity value between the noun information calculated by the second trained model 33C and each patch information is greater than a preset threshold. Specifically, the importance determination unit 111C determines the importance of each of the multiple patch images based on whether the similarity value between the subject information calculated by the second trained model 33C and each patch information is greater than a preset threshold.

[0137] The hardware configuration of the information processing device 100C and the model processing device 30C is the same as that of the information processing device 100 according to Embodiment 1, so a description will be omitted.

[0138] Next, the details of the processing performed by the information processing device 100C will be described with reference to Figures 12 and 15. Figure 15 is a flowchart showing an example of the processing performed by the information processing device 100C according to Embodiment 4. Note that some of the processing performed by the information processing device 100C according to Embodiment 4 is the same as the processing performed by the information processing device 100B according to Embodiment 3, so the same processing as in Embodiment 3 is denoted by the same reference numerals as in Embodiment 3 and its description is omitted.

[0139] As shown in Figure 15, when the information processing device 100C starts processing, it first acquires image information (step ST01). After performing the process in step ST01, the information processing device 100C acquires multiple patch information (step ST02). After performing the process in step ST02, the information processing device 100C acquires request information (step ST03). After performing the process in step ST03, the information processing device 100C generates a subject generation prompt (step ST04). After performing the process in step ST04, the information processing device 100C outputs the subject generation prompt to the first trained model 32B (step ST05).

[0140] When the model processing device 30B receives a subject generation prompt from the information processing device 100C, it inputs the input subject generation prompt to the first trained model 32B, causing the first trained model 32B to extract subject information included in the request information. For example, based on the input of the subject generation prompt, the first trained model 32B extracts subject information indicating the subject "worker" from the request information. Once the subject information has been extracted by the first trained model 32B, the model processing device 30B outputs the extracted subject information to the information processing device 100C.

[0141] When the information processing device 100C performs the processing in step ST05, it obtains subject information from the first trained model 32B (step ST06).

[0142] When the information processing device 100C performs the processing in step ST06, it generates a similarity calculation prompt (step ST07). In this process, the information processing device 100C generates a similarity calculation prompt, which is instruction information that instructs the second trained model 33C to calculate the similarity between the subject information obtained in step ST06 and each patch information obtained in step ST02. For example, if the information processing device 100C has obtained subject information indicating "worker" as the subject information in step ST06, in step ST38 it generates a similarity calculation prompt, which is instruction information indicated by the following string: "Please determine the similarity between the image and "worker" on a scale of 0 to 1. Please output only numerical values."

[0143] When the information processing device 100C performs the processing in step ST07, it outputs a similarity calculation prompt and multiple patch information to the second trained model 33C (step ST38). In this process, the information processing device 100C outputs these similarity calculation prompts and multiple patch information to the second trained model 33C of the model processing device 30 in order to cause the second trained model 33C to generate the similarity between the similarity calculation prompt generated in step ST07 and the multiple patch information obtained in step ST02.

[0144] In step ST38, the model processing device 30C inputs the similarity calculation prompt and multiple patch information received from the information processing device 100C to the second trained model 33C, causing the second trained model 33C to calculate the similarity between the subject information and each patch information. The model processing device 30B outputs the similarity between the subject information and each patch information calculated by the second trained model 33C to the information processing device 100C.

[0145] When the information processing device 100C performs the processing in step ST38, it obtains the similarity between the subject information and each patch information from the second trained model 33C (step ST09). When the information processing device 100C performs the processing in step ST09, it extracts some of the patch information based on the similarity between the subject information and each patch information (step ST10). When the information processing device 100C performs the processing in step ST10, it extracts the noun information contained in the request information (step ST11). When the information processing device 100C performs the processing in step ST11, it converts the noun information into noun features (step ST12). When the information processing device 100C performs the processing in step ST12, it converts the image information and each patch information into features (step ST23).

[0146] When the information processing device 100C performs the processing in step ST23, it calculates the similarity between the noun feature and each patch feature (step ST14). When the information processing device 100C performs the processing in step ST14, it determines the importance of each patch image (step ST25). When the information processing device 100C performs the processing in step ST25, it extracts important features from the image features (step ST16). When the information processing device 100C performs the processing in step ST16, it outputs the image features and important features to the information conversion unit 31 of the model processing device 30B (step ST27). When the information processing device 100C performs the processing in step ST27, it outputs the request information to the first trained model 32B (step ST28). When the information processing device 100C performs the processing in step ST18, it obtains generated information from the first trained model 32B (step ST29). When the information processing device 100C performs the processing in step ST29, it outputs the generated information to the display device 20 (step ST20).

[0147] The information processing device 100C terminates processing after performing the processing in step ST20. For example, the information processing device 100C is configured to perform the processing from step ST01 to step ST20 shown in Figure 15 each time image information is acquired by the image information acquisition unit 103.

[0148] As described above, the information processing device 100C according to Embodiment 4 comprises: an image information acquisition unit 103 that acquires image information representing an image; a patch information acquisition unit 104 that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit 105 that converts the image information, the first patch information and the second patch information into image features, first patch features and second patch features, respectively; an importance determination unit 111C that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; a feature extraction unit 112 that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit 111C; and a request acquisition unit 106 that acquires request information that indicates a user's request regarding an image in natural language. The device is configured to determine the importance of the first patch image and the second patch image based on the first patch features, the second patch features and the request information.

[0149] For example, the information processing device 100C according to Embodiment 4 is configured to generate instruction information for the second trained model 33C to calculate a first similarity between the noun feature and the first patch feature, and a second similarity between the noun feature and the second patch feature, based on noun information, first patch information, and second patch information.

[0150] With this configuration, the information processing device 100C extracts features corresponding to a specific region of the original image from the image features of the original image based on the similarity between the noun information included in the request information and the patch information representing each patch image when the original image is divided into multiple patch images. Therefore, when the trained model generates information based on the extracted features, it becomes possible to have the trained model focus on processing regions of high importance, thereby suppressing hallucination when processing image information with the trained model compared to conventional methods.

[0151] Embodiment 5. Next, the information processing system 1D according to Embodiment 5 will be described with reference to Figures 16 and 17. The information processing system 1D according to Embodiment 5 differs from the information processing system 1 according to Embodiment 1 in some configuration of the information processing device. However, the same configuration as in Embodiment 1 will be given the same names and reference numerals as in Embodiment 1 and will not be described.

[0152] Figure 16 is a block diagram showing the schematic configuration of the information processing system 1D according to Embodiment 5. As shown in Figure 16, the information processing system 1D according to Embodiment 5 comprises an input device 10, a display device 20, a model processing device 30, and an information processing device 100D, which are connected wirelessly or by wire to enable communication between them. The input device 10, the display device 20, the model processing device 30, and the information processing device 100D may also be connected to each other via other devices or communication lines (not shown) to enable information communication between them.

[0153] The information processing device 100D includes an input unit 101, an output unit 102, an image information acquisition unit 103, a patch information acquisition unit 104, an image conversion unit 105, a similarity calculation unit 110D, an importance determination unit 111D, and a feature extraction unit 112D.

[0154] The similarity calculation unit 110D calculates inter-image similarity, which is the similarity between multiple patch features obtained by the image conversion unit 105. In other words, the similarity calculation unit 110D calculates inter-image similarity, which is the similarity between multiple patch features obtained by the image conversion unit 105, including the inter-image similarity between the first patch feature and the second patch feature. For example, the similarity calculation unit 110D calculates the Euclidean distance as the similarity between numerical vectors representing multiple patch features. The similarity calculation unit may also be composed of a large-scale language model that calculates the similarity between multiple input patch features based on the input of multiple patch features.

[0155] Furthermore, the similarity calculation unit 110 may be configured to calculate the similarity between two patch features for all combinations of multiple patch features obtained by the image conversion unit 105, or it may be configured to calculate the similarity between two patch features corresponding to patch images that are adjacent to each other in the original image.

[0156] The importance determination unit 111D determines the importance of each of the multiple patch information obtained by the patch information acquisition unit 104 based on the similarity calculated by the similarity calculation unit 110D. In other words, the importance determination unit 111D determines the importance of each of the multiple patch images, including the first patch image and the second patch image, based on multiple inter-image similarities, including the inter-image similarity between the first patch feature and the second patch feature calculated by the similarity calculation unit 110D.

[0157] For example, the importance determination unit 111D determines whether the patch image corresponding to the inter-image similarity is an important patch image with relatively high importance or an unimportant patch image with relatively low importance, based on the inter-image similarity calculated by the similarity calculation unit 110D. Specifically, the importance determination unit 111D determines whether the patch image corresponding to the inter-image similarity is an important patch image with relatively high importance or an unimportant patch image with relatively low importance, based on whether the inter-image similarity calculated by the similarity calculation unit 110D is higher than a preset threshold. In other words, if the inter-image similarity calculated by the similarity calculation unit 110D is higher than a preset threshold, the importance determination unit 111D determines that the patch image corresponding to the inter-image similarity is an important patch image with relatively high importance, and if the inter-image similarity calculated by the similarity calculation unit 110D is less than or equal to a preset threshold, the importance determination unit 111D determines that the patch image corresponding to the inter-image feature is an unimportant patch image with relatively low importance.

[0158] The feature extraction unit 112D extracts features corresponding to specific regions of the original image from the image features of the original image, based on the judgment result of the importance determination unit 111D. For example, the feature extraction unit 112D extracts features corresponding to regions with higher importance than other regions from the image features of the original image, based on the judgment result of the importance determination unit 111D, as important features. Specifically, the feature extraction unit 112D extracts important features from the image features of the original image, including features corresponding to important patch images, based on the judgment result of the importance determination unit 111D. Details of the process by which the feature extraction unit 112D extracts important features will be described later.

[0159] The hardware configuration of the information processing device 100D is the same as that of the information processing device 100 according to Embodiment 1, so its description will be omitted.

[0160] Next, with reference to Figures 16 and 17, the details of the processing performed by the information processing device 100D will be described. Figure 16 is a flowchart showing an example of the processing performed by the information processing device 100D according to Embodiment 5. Note that some of the processing performed by the information processing device 100D according to Embodiment 5 is the same as the processing performed by the information processing device 100 according to Embodiment 1, so the same processing as in Embodiment 1 is denoted by the same reference numerals as in Embodiment 1 and its description is omitted.

[0161] As shown in Figure 16, when the information processing device 100D starts processing, it first acquires image information (step ST01). After performing the processing in step ST01, the information processing device 100D acquires multiple patch information (step ST02). After performing the processing in step ST03, the information processing device 100D converts the image information and each patch information into feature quantities (step ST13).

[0162] After performing the processing in step ST13, the information processing device 100D calculates the similarity between multiple patch features (step ST24). In this process, the information processing device 100D calculates the inter-image similarity, which is the similarity between multiple patch features obtained in the processing of step ST13.

[0163] When the information processing device 100D performs the processing in step ST24, it determines the importance of each patch image (step ST35). In this process, the information processing device 100D determines the importance of each of the multiple patch images indicated by the multiple patch information obtained in the processing of step ST02, based on the similarity between the multiple patch features calculated in the processing of step ST24.

[0164] When the information processing device 100D performs the processing in step ST15, it extracts important features from the image features (step ST36). In this process, the information processing device 100D extracts features corresponding to a specific region of the original image from the image features of the original image based on the determination result in step ST35. For example, in this process, based on the determination result in step ST36, the information processing device 100D maintains the features corresponding to the important patch image in the image features, removes the features corresponding to one of the two patch images that were deemed unimportant for calculating the similarity between images by replacing the features corresponding to one of the patch images with 0, and replaces the patch features corresponding to the other patch image with the average value of the features of these two patch images, thereby extracting important features, including the features corresponding to the important patch image, from the image features of the original image.

[0165] When the information processing device 100D has completed the processing in step ST16, it outputs important feature quantities to the information conversion unit 31 of the model processing device 30 (step ST17). When the information processing device 100D has completed the processing in step ST17, it terminates its processing. For example, the information processing device 100D is configured to perform the processing from step ST01 to step ST17 shown in Figure 4 each time image information is acquired by the image information acquisition unit 103.

[0166] As described above, the information processing device 100D according to Embodiment 5 comprises: an image information acquisition unit 103 that acquires image information representing an image; a patch information acquisition unit 104 that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit 105 that converts the image information, the first patch information and the second patch information into image features, first patch features and second patch features, respectively; an importance determination unit 111D that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; a feature extraction unit 112D that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit 111D; and a similarity calculation unit 110D that calculates the similarity between the first patch features and the second patch features. The device is configured to determine the importance of the first patch image and the second patch image based on the similarity between the images.

[0167] With this configuration, the information processing device 100D extracts features corresponding to specific regions of the original image from the image features of the original image based on the similarity between each patch image when the original image is divided into multiple patch images. Therefore, when the trained model generates information based on the extracted features, it becomes possible to have the trained model focus on processing regions of high importance, thereby suppressing hallucination when processing image information with the trained model compared to conventional methods.

[0168] In Embodiment 5, the feature extraction unit 112D is configured to extract important features, including those corresponding to important patch images, from the original image's image features by maintaining the features corresponding to important patch images in the image features, deleting the features corresponding to one of the two patch images involved in calculating inter-image similarity that are deemed unimportant patch images by replacing the features corresponding to one of the patch images with 0, and replacing the patch features corresponding to the other patch image with the average value of the features of these two patch images. However, the unit is not limited to this configuration. The feature extraction unit only needs to be configured to extract features corresponding to a specific region of an image from the image features based on the importance of each patch image determined based on inter-image similarity. For example, the feature extraction unit may be configured to delete both features corresponding to the two patch images involved in calculating inter-image similarity that are deemed unimportant patch images from the image features based on the importance of each patch image determined based on inter-image similarity, or it may be configured to delete the features corresponding to one of the two patch images while maintaining the features corresponding to the other patch image.

[0169] In any of the embodiments described above, the information processing device may have some or all of the functions of other devices in the information processing system, some of the functions of the information processing device may be provided in other devices that are communicatively connected to the information processing device, and the devices may be configured so that each function of the information processing device is realized through the cooperation of multiple devices.

[0170] Furthermore, this disclosure allows for free combination of each embodiment, modification of any component of each embodiment, or omission of any component in each embodiment. For example, the information processing device according to embodiments 2 to 4 may include a similarity calculation unit 110D according to embodiment 5 instead of a similarity calculation unit 110, or the information processing device according to embodiment 5 may include a request acquisition unit 106, instruction information generation unit 107, noun information extraction unit 108, string conversion unit 109, and generation information acquisition unit 113 according to embodiment 3.

[0171] The information processing device relating to this disclosure can be used to suppress hallucination when processing image information using a trained model.

[0172] 1 Information processing system, 1A Information processing system, 1B Information processing system, 1C Information processing system, 1D Information processing system, 10 Input device, 20 Display device, 30 Model processing device, 30B Model processing device, 30C Model processing device, 31 Feature conversion unit (information conversion unit), 32 Trained model, 32B First trained model, 33B Second trained model, 33C Second trained model, 100 Information processing device, 100A Information processing device, 100B Information processing device, 100C Information processing device, 100D Information processing device, 100a Processor, 100b Memory, 100c I / O port, 100d Processing circuit, 101 Input unit, 102 Output unit, 103 Image information acquisition unit, 104 Patch information acquisition unit, 105 Image conversion unit, 106 Request acquisition unit, 107 Instruction information generation unit, 107C Instruction information generation unit, 108 Noun information extraction unit, 109 String conversion unit, 110 Similarity calculation unit, 110D Similarity calculation unit, 111 Importance determination unit, 111A Importance determination unit, 111B Importance determination unit, 111C Importance determination unit, 111D Importance determination unit, 112 Feature extraction unit, 112B Feature extraction unit (information extraction unit), 112D Feature extraction unit, 113 Generated information acquisition unit, A1 Area, A2 Area, I1 Original image.

Claims

1. An information processing device comprising: an image information acquisition unit that acquires image information representing an image; a patch information acquisition unit that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit that converts the image information, the first patch information and the second patch information into image features, first patch features and second patch features, respectively; an importance determination unit that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; and a feature extraction unit that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit.

2. The information processing apparatus according to claim 1, further comprising an output unit that outputs the extracted features to a feature conversion unit that converts the extracted features to information that can be processed by a trained model that accepts input of information expressed in natural language.

3. The information processing apparatus according to claim 2, comprising a generation information acquisition unit that acquires generation information generated by the trained model based on the input of the extracted features converted by the feature conversion unit, wherein the output unit outputs the generation information toward the display device so that the generation information is displayed on the display device.

4. The information processing apparatus according to any one of claims 1 to 3, comprising a request acquisition unit that acquires request information indicating a user's request for the image in natural language, wherein the importance determination unit determines the importance of the first patch image and the second patch image based on the first patch feature quantity, the second patch feature quantity and the request information.

5. The information processing apparatus according to claim 4, comprising: a noun information extraction unit that extracts noun information indicating nouns included in the request information by morphological analysis of the request information; and a string conversion unit that converts the noun information into noun features, wherein the importance determination unit determines the importance of the first patch image and the second patch image, respectively, based on the first patch features, the second patch features, and the noun features.

6. The information processing apparatus according to claim 2 or 3, comprising: a request acquisition unit that acquires request information indicating a request from a user regarding the image in natural language; an instruction information generation unit that generates instruction information indicating an instruction for a trained model to generate noun information indicating a noun included in the request information based on the request information; and a string conversion unit that converts the noun information generated by the trained model into noun features, wherein the importance determination unit determines the importance of the first patch image and the second patch image, respectively, based on the first patch features, the second patch features, and the noun features.

7. The information processing apparatus according to claim 6, characterized in that the instruction information generation unit generates instruction information indicating an instruction for the trained model to generate noun information indicating at least one of the subject and object included in the request information, based on the request information.

8. The information processing apparatus according to any one of claims 5 to 7, comprising a similarity calculation unit that calculates a first similarity between the noun feature and the first patch feature, and a second similarity between the noun feature and the second patch feature, wherein the importance determination unit determines the importance of the first patch image and the second patch image based on the first similarity and the second similarity.

9. The information processing apparatus according to any one of claims 1 to 8, characterized in that the feature extraction unit extracts important features of high-importance regions of the image from the image features by deleting features corresponding to one of the first patch image and the second patch image that are included in the image features based on the determination result by the importance determination unit.

10. The information processing apparatus according to claim 6 or 7, characterized in that the instruction information generation unit generates instruction information for causing the trained model to calculate a third similarity between the noun information and the first patch information, and a fourth similarity between the noun information and the second patch information, based on the noun information, the first patch information, and the second patch information.

11. The information processing apparatus according to any one of claims 1 to 7, comprising a similarity calculation unit that calculates the image similarity between the first patch feature quantity and the second patch feature quantity, wherein the importance determination unit determines the importance of the first patch image and the second patch image based on the image similarity.

12. An information processing system characterized by comprising: an information processing device according to claim 2, 3, 6, 7, or 10; and a device having the trained model and the feature conversion unit.

13. A program characterized in that it causes a computer to function as: an image information acquisition unit that acquires image information representing an image; a patch information acquisition unit that acquires first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; an image conversion unit that converts the image information, the first patch information and the second patch information into image features, first patch features and second patch features, respectively; an importance determination unit that determines the importance of the first patch image and the second patch image based on the first patch features and the second patch features; and a feature extraction unit that extracts features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit.

14. An information processing method performed by an apparatus comprising an image information acquisition unit, a patch information acquisition unit, an image conversion unit, an importance determination unit, and a feature extraction unit, the method comprising: the step of the image information acquisition unit acquiring image information representing an image; the step of the patch information acquisition unit acquiring first patch information representing the first patch image and second patch information representing the second patch image when the image is divided into a plurality of patch images including a first patch image and a second patch image based on the image information; the step of the image conversion unit converting the image information, the first patch information, and the second patch information into image features, first patch features, and second patch features, respectively; the step of the importance determination unit determining the importance of the first patch image and the second patch image based on the first patch features and the second patch features; and the step of the feature extraction unit extracting features corresponding to a specific region of the image from the image features based on the determination result by the importance determination unit.