Systems, methods, and apparatus

JP2026125543AActive Publication Date: 2026-08-03WHERE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
WHERE
Filing Date
2025-01-22
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0005】 本発明によれば、2つの画像を比較することで得られる情報を取得することができるシステムを提供できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026125543000001_ABST
    Figure 2026125543000001_ABST
Patent Text Reader

Abstract

The present invention provides a system, method, and apparatus for obtaining information obtained by comparing two images. [Solution] A system comprising a user terminal and a server device capable of communicating with the user terminal, wherein the response acquisition process includes means for sending at least two images and a prompt requesting the language model to output information obtained by comparing the two images as a response from the server device, and means for receiving the information obtained by comparing the two images as a response from the language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system, a method, and an apparatus.

Background Art

[0002] As a technique for comparing two images and detecting differences, a method using an image processing algorithm is known. For example, there is a method in which feature points of each image are extracted, compared, and changes between the images are analyzed.

Summary of the Invention

Problems to be Solved by the Invention

[0003] An object of the present invention is to provide a system capable of acquiring information obtained by comparing two images.

Means for Solving the Problems

[0004] The problems of the present invention are [1] A system including at least one computer device, the system comprising: an image transmission means for transmitting at least two images to a language model; a prompt transmission means for transmitting a prompt asking to output, as an answer, information obtained by comparing the two images, to the language model; and an answer receiving means for receiving, as an answer, information obtained by comparing the two images from the language model; [2] The system according to [1], wherein the two images are images of the same imaging object taken at different times, and the prompt is for determining whether a predetermined event has occurred to the imaging object over time by comparing the two images, and for outputting, as an answer, whether the predetermined event has occurred; [3] The system according to [1] or [2], wherein the two images are images of the same imaging object taken at different times, and the prompt is for identifying an event that has occurred to the imaging object over time by comparing the two images, and for outputting, as an answer, the identified event; [4] A system according to [2] or [3] above, which determines whether a predetermined event has occurred to the subject being photographed over time for each of the multiple image regions constituting the image, and outputs whether or not the predetermined event has occurred as an answer, or which identifies an event that has occurred to the subject being photographed over time for each of the multiple image regions constituting the image, and outputs the identified event as an answer; [5] A system according to any of [1] to [4] above, wherein two images are of the same subject taken at different times, and the prompt includes criteria for determining which category an event that occurred on the subject over time corresponds to by comparing the two images, and requests the system to output as an answer which category the event that occurred on the subject over time corresponds to; [6] A system according to any of [1] to [5] above, wherein the prompt includes criteria for determining whether the two images are subject to evaluation and requests the system to output whether or not they are subject to evaluation as an answer; [7] The system according to any of [1] to [6] above, wherein the prompt includes criteria for determining whether the two images are eligible for evaluation, and requests the system to output whether they are eligible for evaluation as an answer; [8] A system according to any of [1] to [7] above, wherein the two images are of the same subject taken at different times, and the prompt requests that, when comparing the two images, the system output information obtained by comparing the two images based on differences between the two images from a different perspective than differences in color and brightness between the two images that have differed depending on the conditions at the time of shooting; [9] The system according to any of [1] to [8] above, wherein the prompt requests the output of a response in a predetermined output format;

[10] The system described in [9] above, wherein the predetermined output format is JavaScript Object Notation;

[11] The system according to any of [1] to

[10] above, wherein the first of two images is an aerial photograph of the ground, and the second image is an aerial photograph of the ground taken after the first image, and covers the same geographical area as the first image.

[12] A system according to any of

[11] , wherein the first of two images is an aerial photograph of the ground before the disaster occurred, and the second image is an aerial photograph of the ground after the disaster occurred, and the same geographical area as the first image;

[13] The system described in

[12] , wherein the prompt requests the system to determine whether or not damage has occurred in the region by comparing two images and to output as an answer whether or not damage has occurred;

[14] The system according to

[12] or

[13] , wherein the prompt requests the system to identify damage caused by a disaster in the region by comparing two images and to output the identified damage as an answer;

[15] The prompt includes criteria for determining which category the damage caused by the disaster in the region falls into by comparing two images, A system described in any of

[12] to

[14] above, which is required to output as an answer which category the damage caused by a disaster in the aforementioned area falls into;

[16] A method to be performed in a system comprising at least one computer device, comprising: an image transmission step of transmitting at least two images to a language model; a prompt transmission step of transmitting a prompt to the language model requesting that it output information obtained by comparing the two images as an answer; and an answer reception step of receiving information obtained by comparing the two images from the language model as an answer;

[17] An apparatus comprising: an image transmission means for transmitting at least two images to a language model; a prompt transmission means for transmitting a prompt to the language model requesting it to output information obtained by comparing the two images as an answer; and an answer receiving means for receiving information obtained by comparing the two images from the language model as an answer;

[18] A system comprising a user terminal and a server device capable of communicating with the user terminal, wherein the user terminal comprises a first image transmission means for transmitting two images to the server device, and the server device comprises a second image transmission means for transmitting at least two images received from the user terminal to a language model, a prompt transmission means for transmitting a prompt to the language model requesting that information obtained by comparing the two images be output as an answer, a first answer receiving means for receiving information obtained by comparing the two images from the language model as an answer, and an answer transmission means for transmitting the received answer or information obtained from the received answer to the user terminal, and the user terminal comprises a second answer receiving means for receiving the answer or information obtained from the answer from the server device;

[19] A method to be performed in a system comprising a user terminal and a server device capable of communicating with the user terminal, the method comprising: a first image transmission step in which the user terminal transmits two images to the server device; a second image transmission step in which the server device transmits at least two images received from the user terminal to a language model; a prompt transmission step in which the server device transmits a prompt to the language model requesting that information obtained by comparing the two images be output as an answer; a first answer reception step in which the server device receives information obtained by comparing the two images from the language model as an answer; and an answer transmission step in which the server device transmits the received answer or information obtained from the received answer to the user terminal; and a second answer reception step in which the user terminal receives the answer or information obtained from the answer from the server device; This can be resolved. [Effects of the Invention]

[0005] According to the present invention, a system can be provided that can acquire information obtained by comparing two images. [Brief explanation of the drawing]

[0006] [Figure 1] This is a block diagram showing the configuration of the system according to the embodiment. [Figure 2] This is a block diagram showing the hardware configuration of a user terminal according to an embodiment. [Figure 3] This is a block diagram showing the hardware configuration of a server device according to an embodiment. [Figure 4] This diagram shows a flowchart of the response acquisition process 1 according to the embodiment. [Figure 5] This figure shows a flowchart of the response acquisition process 2 according to the embodiment. [Modes for carrying out the invention]

[0007] The following describes embodiments of the present invention, but the present invention is not limited to the following embodiments unless it contradicts the spirit of the invention. The order of each process constituting the flowchart described below is not limited to any order that does not cause contradictions or inconsistencies in the processing content, and it is also possible to omit some of the processes constituting the flowchart or to add new processes to each process constituting the flowchart, as long as it does not contradict the processing content. Furthermore, the device that is the main entity that executes each process constituting the flowchart can be changed to another device unless it contradicts the spirit of the invention. In this case, it is possible to change the processing content so as not to cause contradictions or inconsistencies in the processing content.

[0008] <System> Figure 1 is a block diagram showing the configuration of a system according to an embodiment. System 10 includes at least one computer device. System 10 may also include a user terminal 1, a server device 2, and a language model terminal 3. The user terminal 1, server device 2, and language model 3 are connected to each other via a communication network 4 so that they can communicate with one another. The number of user terminals 1 is not particularly limited, and there may be more than one.

[0009] The system 10 may consist of, for example, one computer device (standalone type), one or more server devices and one or more terminal devices (client-server type), or one host device and one or more guest devices (peer-to-peer type). Furthermore, server device 2 may function distributed across multiple computer devices. For example, instead of server device 2, distributed ledger technology such as blockchain may be used.

[0010] <User terminal> User terminal 1 is a terminal operated by a user of system 10. Here, a user is anyone who uses system 10, and the concept includes operators who run system 10, administrators who manage system 10, and even customers of administrators and operators who use system 10.

[0011] Figure 2 is a block diagram showing the hardware configuration of a user terminal according to an embodiment. User terminal 1 comprises a control unit 11, RAM 12, storage unit 13, input unit 14, display unit 15, and communication interface 16, all of which are connected by a bus.

[0012] The control unit 11 is composed of a CPU and a ROM. The control unit 11 executes a program stored in the storage unit 13 to control the user terminal 1. The RAM 12 is a work area for the control unit 11. The storage unit 13 is a storage medium for storing programs and data. The control unit 11 performs arithmetic processing based on the program and data read from the RAM 12 and the data input at the input unit 14. The program may be stored in a recording medium such as a CD-ROM.

[0013] The display unit 15 has a display screen. The control unit 11 outputs a video signal for displaying an image on the display screen according to the result of the arithmetic processing.

[0014] The communication interface 16 can be connected to the communication network 4 wirelessly or by wire, and can transmit and receive data with other computer devices via the communication network 4. The data received via the communication interface 16 is loaded into the RAM 12, and arithmetic processing is performed by the control unit 11.

[0015] <Server device> FIG. 3 is a block diagram showing the hardware configuration of the server device according to the embodiment. The server device 2 includes at least a control unit 21, a RAM 22, a storage unit 23, and a communication interface 24, which are connected by an internal bus respectively.

[0016] The control unit 21 is composed of a CPU and a ROM, executes a program stored in the storage unit 23, and controls the server device 2. In addition, the control unit 21 includes an internal timer for measuring time. The RAM 22 is a work area for the control unit 21. The storage unit 23 is a storage area for storing programs and data. That is, the storage unit 23 functions as a recording medium storing programs. The control unit 21 reads programs and data from the RAM 22, and performs program execution processing based on the information received from each of the user terminals 1.

[0017] <Language model> Language model 3 is a model that has learned statistical features obtained from a large amount of text data. Based on the received prompt, it can estimate the probability of generating the next word and send the estimated result as a response to the prompt sender. Preferably, language model 3 is a large-scale language model that has been trained on a large dataset and has a vast number of parameters. Furthermore, it is preferable that language model 3 is a multimodal large-scale language model that is capable of image analysis and natural language input and output.

[0018] <Image> Images are not particularly limited, but may be still images or videos. Images may also be a combination of multiple images. The subjects of photography are not particularly limited, but may be living things such as people, animals, and plants; artificial objects such as buildings, roads, facilities, bridges, vehicles, ships, and other items; or natural objects such as rivers, seas, ponds, forests, soil, sand, rocks, and mountains.

[0019] As will be discussed later, the image may be an image of the ground taken from above. Examples of images of the ground taken from above include aerial photographs and satellite images. Aerial photographs are photographs of the ground taken by a camera mounted on an aircraft. Satellite images are image data acquired by sensors mounted on artificial satellites.

[0020] The image data format is not particularly limited. For example, the image data format may be JPEG, BMP, PNG, TIFF, GeoTIFF, or any other format.

[0021] The two images sent to language model 3 may be of the same subject taken at different times. The two images do not necessarily have to be of the same subject taken from the same location or direction, but in order to obtain accurate analysis results from image comparison, it is preferable that they be of the same subject taken from substantially the same location and direction.

[0022] Furthermore, the two images sent to language model 3 may not only be of the same subject taken at different times, but may also be of different subjects. In this case, the two images will be of the same subject taken at different locations.

[0023] For example, in the event of an incident (disaster), even if two man-made structures such as buildings have the same or similar structure (i.e., two man-made structures of the same or similar type), the events that occur and their severity may differ depending on the location of the man-made structures. In such cases, images taken of the two man-made structures can be compared and analyzed.

[0024] For example, in the event of an incident (disaster), even if two man-made structures such as buildings are located in the same or nearby locations or areas, the resulting incident and its severity may differ depending on the structure and type of the structure. In such cases, images of the two man-made structures can be compared and analyzed.

[0025] <prompt> The user operates user terminal 1 to send at least two images to server device 2. Server device 2 generates a prompt that requests information obtained by comparing the images received from user terminal 1 as the answer. Then, it sends the at least two images received from user terminal 1 and the generated prompt to language model 3. Language model 3 generates an answer based on the prompt and sends the generated answer to server device 2. This series of processes corresponds to answer acquisition process 1, which will be described later.

[0026] Furthermore, the user can operate user terminal 1 to directly send two images to language model 3 without going through server device 2, and also send a prompt requesting information obtained by comparing the two images as an answer. Language model 3 generates an answer based on the prompt and sends the generated answer to user terminal 1. This series of processes corresponds to answer acquisition process 2, which will be described later.

[0027] A prompt is textual information that represents the content of an instruction or request made to the language model 3, and may include sentences such as, "Compare the two images and output the information about XX." There are no particular restrictions on the format in which instructions or requests are conveyed in a prompt; in addition to "Please output," other phrases such as "I would like you to output" or "Please tell me" may also be used to convey instructions or requests.

[0028] The content of the prompt is not particularly limited, as long as it asks for the output of information obtained by comparing two images as the answer. Examples of prompt content include the following. Furthermore, prompt examples 1 to 7 shown below may be used individually in one prompt sent to language model 3, or some of these examples may be used in any combination in one prompt sent to language model 3.

[0029] [Prompt Example 1] For example, if two images are of the same subject taken at different times, the prompt may ask the user to compare the two images to determine whether a predetermined event occurred to the subject over time, and to output whether or not the predetermined event occurred as the answer. In this case, the prompt only needs to ask whether or not the predetermined event occurred as the answer, and the format of its expression (i.e., how it is expressed in sentences) is not particularly limited.

[0030] "Whether or not a predetermined event occurred" may mean whether or not a predetermined event occurred between the time the first image was taken and the time the second image was taken, or it may mean whether or not the predetermined event occurred at the time the second image was taken, or it may mean that it includes both.

[0031] The predetermined events can be set as appropriate according to the answers to be obtained using System 10, and are not particularly limited. The predetermined events can be any changes or occurrences that may occur over time. For example, if the subject of photography is a person, animal, or plant, the predetermined events could include the growth of the person, animal, or plant, weight gain or loss, aging, the occurrence of illness or injury (damage), and recovery from illness or injury (damage). If the subject of photography is an artificial object other than a person, animal, or plant, the predetermined events could include damage to the artificial object, recovery from damage, soiling, recovery from soiling, and disappearance.

[0032] The prompt may request that for each of the multiple image regions constituting the image, a predetermined event occur in the subject being photographed over time, and that the prompt output whether or not the predetermined event occurred as an answer. The image transmitted from the user terminal 1 or server device 2 to the language model 3 can be divided into multiple image regions, and for each of these image regions, the prompt can be requested to compare two images and determine whether or not a predetermined event has occurred. For example, by dividing the image into three vertically and three horizontally, the image can be divided into nine image regions, and the prompt can be requested to determine whether or not a predetermined event has occurred in each image region.

[0033] [Prompt Example 2] If two images are of the same subject taken at different times, the prompt may ask the user to compare the two images to identify an event that occurred to the subject over time and output the identified event as the answer. In this case, the prompt only needs to ask for the identification of the event that occurred, and the form of its expression is not particularly limited.

[0034] "Events that occurred" may refer to events that occurred between the time the first image was taken and the time the second image was taken, or events that occurred at the time the second image was taken, or it may include both. Furthermore, "identification of events that occurred" is a concept that includes not only identifying what kind of events occurred, but also identifying the degree to which the events occurred. If the subject of the photograph is a person, animal, or plant, events can be identified as growth, weight gain / loss, aging, occurrence of illness / injury (damage), recovery from illness / injury (damage), etc. If the subject of the photograph is an artificial object other than a person or animal, events can be identified as damage to the artificial object, recovery from damage, soiling, recovery from soiling, disappearance, etc.

[0035] The prompt may request that the user identify an event that occurred in the subject being photographed over time for each of the multiple image regions that make up the image, and output the identified event as the answer. The image transmitted from the user terminal 1 or server device 2 to the language model 3 can be divided into multiple image regions, and for each of these image regions, the user can be asked to identify an event that occurred by comparing two images. For example, by dividing the image into three parts vertically and three parts horizontally, the image can be divided into nine image regions, and the user can be asked to identify an event that occurred in each image region.

[0036] [Prompt Example 3] Furthermore, for example, if two images are of the same subject taken at different times, the prompt may include criteria for identifying which category the events that occurred on the subject over time belong to by comparing the two images, and may request the system to output which category the events that occurred on the subject belong to as an answer. The format of the prompt is not particularly limited, as long as it includes criteria for identifying which category the events that occurred belong to and requests the system to output which category the events belong to as an answer.

[0037] "Events that occurred" may refer to events that occurred between the time the first image was taken and the time the second image was taken, or events that occurred at the time the second image was taken, or it may include both. Furthermore, "identification of events that occurred" is a concept that includes not only identifying what kind of events occurred, but also identifying the degree to which the events occurred. If the subject of the photograph is a person, animal, or plant, events can be identified as growth, weight gain / loss, aging, occurrence of illness / injury (damage), recovery from illness / injury (damage), etc. If the subject of the photograph is an artificial object other than a person or animal, events can be identified as damage to the artificial object, recovery from damage, soiling, recovery from soiling, disappearance, etc.

[0038] A "classification" is used to categorize the events that have occurred, and may be a tiered classification or a classification of different types. A tiered classification means that the degree of the event changes in stages, such as from level 1 to level 5, where the degree of the event decreases or increases as the level number decreases. For example, if the event is damage to a building, the prompt may be written so that the degree of damage increases as the level number increases, from level 1 to level 5.

[0039] In this case, it is preferable that the prompt includes a criterion for determining the severity of the event that occurred. For example, the prompt can be written to determine the severity of the event based on the area of ​​the region in the image that has changed, after comparing two images. For example, for each of the five levels from Level 1 to Level 5, a range of reference ratios of the area of ​​the region in the image that has changed relative to the entire image is described. Alternatively, the prompt can be written to determine the severity of the event based on the magnitude of the change in the color of each pixel in the image (e.g., RGB values), after comparing two images.

[0040] [Prompt Example 4] The prompt may include criteria for determining whether two images are subject to evaluation and may request the output of whether or not they are subject to evaluation as the answer. The format of the prompt is not particularly limited, as long as it includes criteria for determining whether two images are subject to evaluation and outputs whether or not they are subject to evaluation as the answer.

[0041] In determining whether or not something is subject to evaluation, the ratio of the area occupied by the subject to evaluation to the total area of ​​the image can be used as a criterion. For example, when determining whether or not there is damage on land in the event of a disaster, or when identifying the type and extent of damage, the sea is not subject to evaluation. Therefore, for each of the two images, if the area of ​​pixels corresponding to the sea, which is not subject to evaluation, exceeds a predetermined ratio, it can be requested to output that it is not subject to evaluation.

[0042] [Prompt Example 5] The prompt may include criteria for determining whether two images are evaluable and may request that the response be whether or not they are evaluable. The format of the prompt is not particularly limited, as long as it includes criteria for determining whether two images are evaluable and requests that the response be whether or not they are evaluable.

[0043] For example, if at least one of the two images is of poor quality and it is unclear what the image represents, it may be difficult to accurately identify the subject and evaluate it. The prompt may include criteria related to image quality, such as image resolution and noise level, to determine whether the two images are suitable for evaluation.

[0044] [Prompt Example 6] When two images are of the same subject taken at different times, the prompt may ask the user to compare the two images and output information obtained by comparing the two images based on differences in a different perspective than differences in color and brightness between the two images that may have occurred due to the conditions at the time of shooting. The prompt only needs to ask the user to make a judgment by comparing the images without considering differences in color and brightness caused by sunlight conditions, weather, time of day, etc., and to focus on substantial changes (shape of the subject, damaged areas), and the form of its expression is not particularly limited.

[0045] [Prompt Example 7] A prompt may request that the response be output according to a predetermined output format. The prompt only needs to define how the response format should be expressed, and the format of that expression is not particularly limited.

[0046] The specified output format is not particularly limited, but examples include JavaScript Object Notation (JSON) format, XML format, and CSV format. If the prompt states "Output the answer in JSON format," then structured data consisting of key-value pairs can be received as the answer from language model 3. This makes it easier for user terminal 1 or server device 2 to parse the answer and to cooperate with other programs and systems.

[0047] <Response Acquisition Process 1> The following describes the response acquisition process 1 in System 10. Figure 4 is a flowchart of the response acquisition process according to the embodiment. Here, we mainly describe the case in which two images are compared: an image of the ground taken from above before the disaster occurred, and an image of the ground taken from above after some time has passed and the disaster occurred, and information about the damage caused by the disaster in the real world corresponding to these images is acquired using the language model 3.

[0048] First, the user logs in to system 10 by operating user terminal 1 (step S1). When logging in, the user enters their user ID and password on user terminal 1, and authentication is performed based on this information. Access to system 10 from user terminal 1 may be done using either a browser application or a native application.

[0049] The user operates user terminal 1 to send images to server device 2 of the same geographical area (i.e., the same region) taken from above before the disaster (e.g., satellite imagery) and images taken from above after the disaster (e.g., satellite imagery) to the same geographical area (region) taken before the disaster and the geographical area (region) taken after the disaster may be exactly the same, but it is sufficient if there is at least some overlap, and they do not have to be exactly the same.

[0050] Server device 2 receives two images from user terminal 1 (step S3). In step S2, in addition to sending image files from user terminal 1 to server device 2, text information entered at user terminal 1 may also be sent. User terminal 1 can input conditions and rules for determining the two images as text information and send this to server device 2.

[0051] Next, the server device 2 generates a prompt in the control unit 21 (step S4). If text information about conditions and rules is received in step S3 from the user terminal 1, this text information can be used to generate a prompt. The server device 2 generates a prompt to request the language model 3 to output information obtained by comparing two satellite images, one before the disaster and one after the disaster. Specifically, the prompt is written to request the language model 3 to determine whether there is damage, the level of damage, whether there is damage to roads and buildings, whether the image is analyzable, and whether the image is not excluded from evaluation. Furthermore, the prompt can be written to request explanations regarding the degree of damage to each area in the image and the urgency of the response to each area.

[0052] The following are examples of specific prompts: "These images show the situation before and after the disaster. Compare the two images to analyze the extent of the damage caused by the disaster, and then answer the 'questions' according to the 'rules'." #Answer items If damage has occurred in the area corresponding to the image, set "is_disastrous" to "true". If no damage has occurred, set it to "false". Based on the area of ​​the damaged buildings and infrastructure (public facilities such as roads, water, electricity, and gas), please set the level of damage (disastrous_level) according to the following criteria: Level 1: Buildings and infrastructure suffer little to no damage. Level 2: Damage has occurred in less than 10% of the image. Level 3: Damage has occurred in 10-30% of the image. Level 4: Damage has occurred in 30-50% of the image. Level 5: Damage has occurred in more than 50% of the image. Identify any damage to roads or buildings. If damage is found, set "road_damage" and "building_damage" to "true". If there is no damage, set them to "false". If more than 90% of the image area is "sea," it will not be evaluated. In that case, please set "is_sea" to "true." If more than 10% of the image area is "land," it will be evaluated. In that case, please set "is_sea" to "false." If the image is unclear and cannot be analyzed, set "uninterpretable" to "true". If the image can be analyzed, set it to "false". • Divide the image into multiple areas and describe the degree of damage to buildings and infrastructure in each area (e.g., upper area, middle area, lower area, etc.) and the urgency of the response in the "response" section. #rule Please output your response in JSON format. • The overall color and brightness of an image may vary depending on the shooting environment. Please disregard these differences and focus your judgment on the differences related to the damage caused by the disaster.

[0053] The method for generating a prompt is not particularly limited, but for example, a prompt can be generated based on a predefined phrase stored in the server device 2. If text information about conditions or rules is received in step S3, the predefined phrase can be edited using the text information to generate a prompt.

[0054] When a prompt is generated, the server device 2 sends the two images and the prompt to the language model 3 (step S5). When the language model 3 receives the two images and the prompt (step S6), the language model 3 generates a response based on the two images and the prompt it received (step S7). The language model 3 compares the two images and determines or identifies whether there is damage, the level of damage, the condition of roads and buildings, whether the images are subject to evaluation, whether image analysis is possible, the urgency of the response, etc., and generates a response.

[0055] Next, the generated response is sent from the language model 3 to the server device 2 (step S8), and the server device 2 receives it (step S9). The server device 2 stores the transmitted image, prompt, and the response received from the language model 3 in the storage unit 23 (step S10). Next, the server device 2 sends the response received from the language model 3, or information newly generated based on that response, to the user terminal 1 (step S11).

[0056] When user terminal 1 receives a response (step S12), user terminal 1 generates a display image in a predetermined format based on the received response and displays the generated display image on the display screen (step S13). Steps S1 to S13 complete the response acquisition process 1.

[0057] Furthermore, in step S9, when the server device 2 receives a response, it is possible to automatically configure the system so that the received response or newly generated information based on the response can be viewed on a user terminal other than user terminal 1, or this can be configured manually.

[0058] In step S13, based on the response received from the server device 2 in step S12, new information and / or display images are automatically generated according to a predetermined program, and the generated information and / or display images can be displayed on the display screen of the user terminal 1.

[0059] Furthermore, in step S11, based on the response received from the language model 3 in step S9, new information and / or display images are automatically generated according to a predetermined program, and the generated information and / or display images are sent to the user terminal 1, and can be displayed on the display screen of the user terminal 1 in step S13.

[0060] Furthermore, the information contained in the response received in step S9 or S12, and / or the information and / or the automatically generated display images based on the response received in step S9 or S12 according to a predetermined program, can be uploaded to server device 2 or another server device different from server device 2, and made viewable on the display screen of user terminal 1 and / or other user terminals operated by other users different from the user.

[0061] Furthermore, the information contained in the response received in step S9 or S12, and / or the information and / or displayed images processed and edited on user terminal 1 or another user terminal different from user terminal 1 based on the response received in step S9 or S12, can be uploaded to server device 2 or another server device different from server device 2, and made viewable on the display screen of user terminal 1 and / or other user terminals operated by other users different from the user.

[0062] For example, in steps S1 to S7, two images are compared within a single geographical area to obtain information such as whether or not there is damage, the level of damage, the condition of roads and buildings, whether the image is subject to evaluation, whether or not image analysis is possible, and the urgency of the response. This process in steps S1 to S7 is repeated for multiple geographical areas. Then, for each of the multiple geographical areas obtained, the information and / or displayed images, which have been processed and edited automatically or manually based on these answers, are uploaded to server device 2 or another server device different from server device 2, and can be viewed on the display screen of user terminal 1 and / or other user terminals operated by other users different from the user.

[0063] <Response Acquisition Process 2> The following describes the response acquisition process 2 in System 10. Figure 5 is a flowchart of the response acquisition process according to the embodiment. Here, we mainly describe the case where two images are compared: an image of the ground taken from above before the disaster occurred, and an image of the ground taken from above after some time has passed and the disaster occurred, and information about the damage caused by the disaster in the real world corresponding to these images is acquired using the language model 3.

[0064] First, the user logs in to system 10 by operating user terminal 1 (step S21). When logging in, the user enters their user ID and password on user terminal 1, and authentication is performed based on this information. Access to system 10 from user terminal 1 may be done using either a browser application or a native application.

[0065] The user operates user terminal 1 to send satellite images of the same region (same geographical area) before and after the disaster, as well as a prompt, to language model 3 (step S22). The prompt sent to language model 3 requests that it output information obtained by comparing the two satellite images, one before and one after the disaster. Specifically, the prompt is written to ask language model 3 to determine whether there is damage, the level of damage, whether there is damage to roads or buildings, whether the image is analyzable, and whether the image is ineligible for evaluation. Furthermore, the prompt can also be written to request explanations regarding the degree of damage in each area within the image and the urgency of the response for each area. Specific examples of prompts can be the same as those exemplified above.

[0066] In step S22, the user may operate user terminal 1 to input prompts each time, reuse a prompt previously saved on user terminal 1 and send it to language model 3, or reuse a saved prompt while making edits before sending it to language model 3.

[0067] When language model 3 receives two images and a prompt (step S23), language model 3 generates a response based on the two received images and prompt (step S24). Language model 3 compares the two images and determines or identifies whether there is damage, the level of damage, the condition of roads and buildings, whether the image is subject to evaluation, whether image analysis is possible, the urgency of the response, etc., and generates a response.

[0068] Next, the generated response is sent from the language model 3 to the user terminal 1 (step S25), and the user terminal 1 receives it (step S26). Upon receiving the response from the language model 3, the user terminal 1 generates a display image in a predetermined format based on the received response or information newly generated based on the response, and displays the generated display image on the display screen (step S27). Steps S21 to S27 complete the response acquisition process 2.

[0069] In step S27, based on the response received from language model 3 in step S26, new information and / or display images are automatically generated according to a predetermined program, and the generated information and / or display images can be displayed on the display screen of user terminal 1.

[0070] Furthermore, the information contained in the response received in step S26, and / or the information and / or the automatically generated display images based on the response received in step S26 according to a predetermined program, can be uploaded to server device 2 or another server device different from server device 2, and made viewable on the display screen of user terminal 1 and / or other user terminals operated by other users different from the user.

[0071] Furthermore, the information contained in the response received in step S26, and / or information and / or displayed images processed and edited on user terminal 1 or another user terminal different from user terminal 1 based on the response received in step S26, can be uploaded to server device 2 or another server device different from server device 2, and made viewable on the display screen of user terminal 1 and / or other user terminals operated by other users different from the user.

[0072] For example, in steps S21 to S26, two images are compared for a single geographical area to obtain information such as whether or not there is damage, the level of damage, the condition of roads and buildings, whether the image is subject to evaluation, whether or not image analysis is possible, and the urgency of the response. This process in steps S21 to S26 is repeated for multiple geographical areas. For each of the multiple geographical areas obtained, the information and / or displayed images, including whether or not there is damage, the level of damage, the condition of roads and buildings, whether or not the image is subject to evaluation, whether or not image analysis is possible, and the urgency of the response, or information processed and edited automatically or manually based on these answers, can be uploaded to server device 2 or another server device different from server device 2, and made viewable on the display screen of user terminal 1 and / or other user terminals operated by other users different from the user. In this case, user terminal 1 or other user terminals can operate to select which of the multiple geographical areas to display information such as whether or not there is damage, the level of damage, the condition of roads and buildings, whether or not the image is subject to evaluation, whether or not image analysis is possible, and the urgency of the response.

[0073] According to the present invention, information obtained by comparing two images can be acquired by sending two images to a language model and sending a prompt to the language model that requests it to output information obtained by comparing the two images as an answer.

[0074] The two images are of the same subject, taken at different times. The prompt asks the user to compare the two images to determine whether or not a predetermined event occurred with respect to the subject. This allows the user to obtain an answer regarding whether or not a predetermined event occurred with respect to the subject.

[0075] The two images are of the same subject, taken at different times. The prompt asks the user to compare the two images to identify the event that occurred on the subject, thereby obtaining the event that occurred on the subject as the answer.

[0076] By requiring the determination of whether a predetermined event occurred in each of the multiple image regions that make up the image, it is possible to obtain an answer regarding whether a predetermined event occurred in each of the multiple image regions that make up the image, based on two images. By requiring the identification of the event that occurred in each of the multiple image regions that make up the image, it is possible to obtain the event that occurred in each of the multiple image regions that make up the image as an answer.

[0077] The two images are of the same subject, taken at different times. The prompt includes criteria for identifying which category the event that occurred on the subject falls into by comparing the two images, and asks the user to identify which category the event that occurred on the subject falls into, thereby allowing the user to obtain an answer regarding which category the event that occurred on the subject falls into.

[0078] The prompt includes criteria for determining whether the two images are subject to evaluation, and asks the user to determine whether they are subject to evaluation, thereby allowing the user to obtain a response regarding whether the two images are subject to evaluation.

[0079] The prompt includes criteria for determining whether the two images are evaluable, and asks the user to determine whether they are evaluable, thereby providing an answer as to whether the two images are evaluable.

[0080] The two images are of the same subject, taken at different times. The prompt asks the user to output information obtained by comparing the two images based on differences in color and brightness that arise from the conditions at the time of shooting, rather than differences in color and brightness that occur depending on the shooting conditions. This makes it easier to obtain answers based on essential differences and changes that differ from differences in color and brightness that occur depending on the shooting conditions.

[0081] The prompt requests the user to output a response according to a predetermined output format, which facilitates subsequent processing and integration with other applications. Preferably, the predetermined output format is JavaScript Object Notation.

[0082] If the first of the two images is an aerial photograph of the ground, and the second image is an aerial photograph of the ground taken after the first image, but covering the same geographical area as the first image, then information about changes in this area can be obtained as a response based on the aerial photograph of the ground. If the first of the two images is an aerial photograph of the ground taken before the disaster occurred, and the second image is an aerial photograph of the ground taken after the disaster occurred, but covering the same geographical area as the first image, then information about damage caused by the disaster can be obtained as a response.

[0083] The prompt asks the user to compare two images to determine whether or not damage from a disaster has occurred, thereby allowing the system to obtain an answer regarding whether or not damage from a disaster has occurred.

[0084] The prompt asks the user to identify the damage caused by the disaster by comparing two images, thereby allowing the user to obtain the disaster as the answer.

[0085] The prompt includes criteria for identifying which category the damage caused by the disaster falls into by comparing two images, and asks the user to identify which category the damage caused by the disaster falls into, thereby allowing the user to obtain an answer regarding which category the disaster falls into.

[0086] 1. User terminal, 2. Server device, 3. Language model 4. Communication Network 10 Systems 11 Control unit, 12 RAM, 13 Storage unit 14 Input section, 15 Display section, 16 Communication interface 21 Control unit, 22 RAM, 23 Storage unit 24 Communication Interfaces

Claims

1. A system comprising at least one computer device, An image transmission means for sending at least two images to a language model, A prompt sending means sends a prompt to a language model that requests the model to output information obtained by comparing two images as an answer, A response receiving means that receives information obtained by comparing two images from a language model as a response, and A system equipped with these features.

2. The two images are of the same subject, taken at different times. The aforementioned prompt, The system according to claim 1, wherein by comparing two images, it determines whether or not a predetermined event occurred to the subject being photographed over time, and outputs whether or not the predetermined event occurred as an answer.

3. The two images are of the same subject, taken at different times. The aforementioned prompt, The system according to claim 1, which compares two images to identify an event that occurred on the subject being photographed over time, and outputs the identified event as the answer.

4. The system is required to determine whether a predetermined event occurred in the subject being photographed over time for each of the multiple image regions that make up the image, and to output whether or not the predetermined event occurred as an answer, or The system according to claim 2 or 3, wherein for each of the multiple image regions constituting the image, it identifies an event that occurred on the subject being photographed over time, and outputs the identified event as an answer.

5. The two images are of the same subject, taken at different times. The aforementioned prompt, The system includes criteria for determining which category an event that occurred to the subject over time falls into by comparing two images. The system according to claim 1, which requires outputting as an answer which category the events that occurred to the subject being photographed over time belong to.

6. The aforementioned prompt, The system according to claim 1 or 2, which includes criteria for determining whether two images are subject to evaluation, and is required to output whether or not they are subject to evaluation as an answer.

7. The aforementioned prompt, The system according to claim 1 or 2, which includes criteria for determining whether two images are eligible for evaluation, and is required to output whether or not they are eligible for evaluation as an answer.

8. The two images are of the same subject, taken at different times. The aforementioned prompt, The system according to claim 1 or 2, wherein when comparing two images, the system is required to output information obtained by comparing the two images based on differences between the two images from different perspectives, including differences in color and brightness between the two images that may have differed depending on the shooting conditions.

9. The aforementioned prompt, The system according to claim 1 or 2, which requires outputting a response in a predetermined output format.

10. The system according to claim 9, wherein the predetermined output format is a JavaScript Object Notation.

11. The system according to claim 1, wherein the first of the two images is an aerial photograph of the ground, and the second image is an aerial photograph of the ground taken after the first image, and captures the same geographical area as the first image.

12. The system according to claim 11, wherein the first of the two images is an aerial photograph of the ground before the disaster occurred, and the second image is an aerial photograph of the ground after the disaster occurred, and the same geographical area as the first image.

13. The aforementioned prompt, The system according to claim 12, which determines whether or not damage has occurred in the region by comparing two images and outputs whether or not damage has occurred as an answer.

14. The aforementioned prompt, The system according to claim 12, which compares two images to identify damage caused by a disaster in the aforementioned region and outputs the identified damage as an answer.

15. The aforementioned prompt, By comparing two images, the system includes criteria for determining which category the damage caused by the disaster in the aforementioned region falls into. The system according to claim 12, which requires outputting as an answer which category the damage caused by a disaster occurring in the aforementioned area falls into.

16. A method performed in a system comprising at least one computer device, An image transmission step that sends at least two images to a language model, A prompt sending step sends a prompt to a language model that requests it to output information obtained by comparing two images as an answer, The response receiving step involves receiving information obtained by comparing two images from a language model as the response, and A method having

17. An image transmission means for sending at least two images to a language model, A prompt sending means sends a prompt to a language model that requests the model to output information obtained by comparing two images as an answer, A response receiving means that receives information obtained by comparing two images from a language model as a response, and A device equipped with the following features.

18. A system comprising a user terminal and a server device capable of communicating with the user terminal, The user terminal, First image transmission means for transmitting two images to a server device. Equipped with, The server device A second image transmission means that transmits at least two images received from a user terminal to a language model, A prompt sending means sends a prompt to a language model that requests the model to output information obtained by comparing two images as an answer, A first response receiving means receives information obtained by comparing two images from a language model as a response, A response transmission means that transmits the received response or information obtained from the received response to the user terminal, Equipped with, The user terminal, Second response receiving means for receiving a response or information obtained from a response from a server device. A system equipped with these features.

19. A method to be performed in a system comprising a user terminal and a server device capable of communicating with the user terminal, The user terminal, First image transmission step: Sending two images to the server device. It has, The server device A second image transmission step involves sending at least two images received from the user terminal to a language model, A prompt sending step sends a prompt to a language model that requests it to output information obtained by comparing two images as an answer, The first response receiving step involves receiving a response from the language model, A response transmission step that sends the received response or information obtained from the received response to the user terminal. It has, The user terminal, Second response receiving step: Receiving a response or information obtained from the response from the server device. A method having