Image recognition systems, methods, and programs

The image recognition system addresses the challenge of temporal recognition by combining images with chronological order and using prompts to enhance generative AI's understanding of temporal changes, ensuring accurate and consistent responses.

JP2026121262AActive Publication Date: 2026-07-23SOFTCREATE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTCREATE CORP
Filing Date
2025-08-12
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Generative AI systems struggle to understand temporal changes and causal relationships in image data due to processing limitations, leading to difficulties in recognizing objects that change over time, as they treat each image as independent information without considering the chronological order.

Method used

An image recognition system that combines multiple images taken at different times into a single data set with chronological order, using predetermined prompts and constraints to ask generative AI questions, integrating responses, and managing data in user-selected environments.

Benefits of technology

Enables accurate recognition of objects changing over time by understanding temporal patterns and reducing ambiguity, improving visual quality, and ensuring consistent responses through noise removal and resolution adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026121262000001_ABST
    Figure 2026121262000001_ABST
Patent Text Reader

Abstract

By assigning sequential numbers to two or more images and combining them into a single image, and by including the sequential order in the prompt, the generating AI can recognize objects that change over time in those two or more images. [Solution] The image recognition system uses still image data, which consists of two or more images of an object taken at different times and combined into a single image data, and which is numbered to indicate the temporal order in which the two or more images included in the still image data were taken, along with a predetermined prompt, to ask a predetermined question about the situation to a generating AI and obtain an answer from the generating AI.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image recognition system, an image recognition method, and a program for recognizing an object that changes over time.

Background Art

[0002] In recent image recognition technologies, by incorporating generative AI as part of multimodal AI, specific objects and patterns are identified from images. Specifically, multimodal AI such as LLM (Large Language Model), which has the ability to integrate and process multiple data formats such as text, images, and voices, complements generative AI that generates new data based on input data, and analyzes images to generate their descriptions.

[0003] For example, in Non-Patent Document 1, text output (natural language, code, etc.) is generated for an input in which text and images are mixed. Also, for example, in Non-Patent Document 2, information that deepens the user's understanding is provided about what the user is looking at in the surrounding environment of the pass-through.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

[0005] However, because generative AI processes input data as individual static pieces of information, it cannot directly understand changes or causal relationships along a time axis. For example, even when given a series of images, it cannot recognize them as a "temporal flow," and each image is treated as independent information.

[0006] Furthermore, there is an upper limit to the amount of information that generating AI can process, and it cannot take in vast amounts of data at once. As a result, it is difficult to directly import videos and continuous image data, and it is necessary to thin out or summarize the data as needed.

[0007] Due to these constraints, when generating AI processes continuous image data, it is generally necessary to divide the data into time slices (still images divided by time) and read each slice individually. While this allows for recognition of objects and features contained in individual images (image recognition), the lack of a concept of time prevents the AI ​​from understanding the continuity or causal relationships between images.

[0008] For example, even if we recognize that "a ball has been thrown" in one image, it is extremely difficult to connect that to the fact that "the ball is moving" in the next image and grasp the entire situation (situational recognition) that "the ball is moving as a result of being thrown." Therefore, to achieve advanced situational recognition that takes the time axis into account, it is necessary to combine generative AI with a system specialized in temporal data processing, rather than using generative AI alone.

[0009] In view of these problems, the present invention aims to provide an image recognition system, an image recognition method, and a program that enable a generating AI to recognize objects that change chronologically in the two or more images by assigning numbers indicating the chronological order to two or more images and combining them into a single image, and by including the chronological order in the prompt. [Means for solving the problem]

[0010] The present invention provides the following solutions.

[0011] According to the invention relating to the first feature, An image recognition system that recognizes the state of an object that changes over time, The present invention provides an image recognition system comprising still image data, which consists of two or more images of the aforementioned object taken at different times and combined into a single image data, and which is numbered to indicate the temporal order in which the two or more images included in the still image data were taken, and a predetermined prompt, which is used to ask a predetermined question to a generating AI about the situation and to obtain an answer from the generating AI.

[0012] According to the invention relating to the first feature, by combining multiple images taken at different times and assigning numbers to the captured images to indicate their temporal order, it becomes possible to provide a foundation for a generative AI to accurately understand the situation. By asking the generative AI questions using these multiple images and predetermined prompts, the AI ​​can detect changes and patterns that humans tend to overlook.

[0013] The invention relating to the second feature is the invention relating to the first feature, The aforementioned predetermined prompt includes a role for responding, predetermined constraints, and an output format. We provide an image recognition system.

[0014] According to the invention relating to the second feature, by including a role in the prompt given to the generating AI, the position and perspective from which the generating AI responds becomes clear, the direction of the response is determined, and it becomes possible to prevent ambiguous answers and inappropriate content. Furthermore, by including constraints in the prompt, it becomes possible to eliminate unnecessary information and content that may cause misunderstandings. In addition, by specifying the output format in the prompt, consistency in the response is maintained, and the user can obtain information in a format that is easy to use directly.

[0015] The invention according to the third feature is the invention according to the first feature, where the still image data is extracted at a predetermined interval from a moving image of the object, and provides an image recognition system.

[0016] According to the invention according to the third feature, by using the still images extracted at a predetermined interval as the still images given to the generation AI, it becomes easier for the generation AI to analyze temporal changes and enables easy tracking.

[0017] The invention according to the fourth feature is the invention according to the second feature, where the predetermined constraints further include the obtained answer, and provides an image recognition system.

[0018] According to the invention according to the fourth feature, by generating new information based on the inference as the answer obtained immediately before, the analysis flow is not interrupted and consistent results can be obtained, so that a consistent answer can be obtained.

[0019] The invention according to the fifth feature is the invention according to the first feature, where noise removal, resolution adjustment, or image normalization is performed on the two or more images before generating the still image data, and provides an image recognition system.

[0020] According to the invention according to the fifth feature, by removing unnecessary noise included in the image, the visual quality is improved and the accuracy in generating still image data is enhanced. Also, by adjusting the resolution, it becomes possible to optimally adjust the size of the still image data to be generated while maintaining the details of the image. Further, by normalizing the image, it becomes possible to more accurately recognize the differences between images by unifying the brightness and contrast of the images.

[0021] The invention according to the sixth feature is the invention according to the first feature, The question section asks the plurality of generative AIs the predetermined questions regarding the situation, obtains the answers from the plurality of generative AIs, and integrates them. Provided is an image recognition system.

[0022] According to the invention according to the sixth feature, by integrating the answers from the plurality of generative AIs, it is possible to compensate for the weaknesses of individual AIs and obtain more accurate and reliable answers.

[0023] The invention according to the seventh feature is the invention according to the first feature, When voice data, sensor data, or text data is associated with each of the two or more images, the voice data, the sensor data, or the text data is integrated into the still image data. Provided is an image recognition system. [[ID=…]] [[ID=…]]

[0024] [[ID=…]] According to the invention according to the seventh feature, by adding text data such as voice data or explanatory text related to the image, the intention and content of the image become clear. Also, by providing environmental information (temperature, humidity, location information, etc.) based on sensor data related to the image, the background and situation of the image become clearer in more detail. As a result, more accurate judgment and prediction by the generative AI become possible.

[0025] The invention according to the eighth feature is the invention according to the first feature, The question section outputs the still image data and the answers of the generative AI, and receives the predetermined questions and feedback from the user. Provided is an image recognition system.

[0026] According to the invention according to the eighth feature, while outputting the still image data and the answers of the generative AI, by receiving questions and feedback from the user, the user can not only obtain the necessary information but also reflect their opinions and doubts in the questions and feedback to the generative AI.

[0027] <^ The invention relating to the ninth feature is the invention relating to the eighth feature, Based on the received response or feedback, the selection of two or more images is made, and the still image data is generated. We provide an image recognition system.

[0028] According to the invention relating to the ninth feature, by outputting still image data and the generating AI's response while accepting questions and feedback from the user, the user can not only obtain the necessary information but also reflect their own opinions and questions in the questions and feedback to the generating AI.

[0029] An invention relating to the tenth feature is an invention relating to any of the first to ninth features, It runs in either an on-premises computing environment or a cloud computing environment, and the static image data and the AI's response are managed in an environment desired by the user. We provide an image recognition system.

[0030] According to the invention relating to the tenth feature, for example, it becomes possible to manage data by selecting an environment that suits the user's needs, such as selecting an on-premises computing environment when security requirements are strict, selecting a cloud computing environment when scalability is required, or performing specific processing in an on-premises computer environment and other processing in a cloud computing environment.

[0031] Although this invention falls under the category of computer systems, it exhibits similar functions and effects in other categories such as methods and programs, depending on the category. [Effects of the Invention]

[0032] According to the present invention, by assigning numbers indicating the chronological order to two or more images and combining them into a single image, and by including the chronological order in the prompt, it is possible to provide an image recognition system, image recognition method, and program that enables the generating AI to recognize objects that change chronologically in the two or more images. [Brief explanation of the drawing]

[0033] [Figure 1] This is a diagram illustrating an overview of Image Recognition System 1, which is one embodiment of the present invention. [Figure 2] This is a diagram showing the configuration of the image recognition system 1 of this embodiment. [Figure 3] This is a flowchart of the image recognition process performed by the image recognition system 1 of this embodiment. [Figure 4] This figure shows an example of two or more images 100 of an object taken at different times in the image recognition system 1 of this embodiment. [Figure 5] This figure shows an example of a single still image data 200 formed by combining images in the image recognition system 1 of this embodiment. [Figure 6] This figure shows an example in which a temporal order is assigned to each image in the still image data 200 in the image recognition system 1 of this embodiment. [Figure 7] This figure shows another example in the image recognition system 1 of this embodiment, in which a temporal order is assigned to each image in the still image data 200. [Modes for carrying out the invention]

[0034] The best mode for carrying out the present invention will be described below with reference to the figures. However, this is merely one example, and the technical scope of the present invention is not limited thereto.

[0035] [Overview of Image Recognition System 1] An overview of an image recognition system 1, which is one embodiment of the present invention, will be described based on Figure 1. Figure 1 is a diagram illustrating the overview of an image recognition system 1, which is one embodiment of the present invention. The image recognition system 1 consists of a computer 2 and is a computer system for recognizing the state of an object that changes over time.

[0036] The computer 2 of the image recognition system 1 may be, for example, a desktop computer, a laptop computer, a server, a mobile device such as a smartphone or tablet, or a wearable device such as a head-mounted display like smart glasses or a smartwatch.

[0037] Furthermore, the computer 2 of the image recognition system 1 may be implemented as, for example, a single terminal device, multiple terminal devices, or a virtual device such as a cloud computer. Alternatively, it may be implemented in an on-premises environment using a dedicated server and network.

[0038] Furthermore, the image recognition system 1 may consist of the terminal device described above instead of the computer 2.

[0039] The computer 2 of the image recognition system 1 is connected to the aforementioned terminal device, other terminals and devices, etc., via a public network, etc., enabling data communication, and performs the sending and receiving of necessary data and information.

[0040] Next, we will describe the overview of the processes performed by the image recognition system 1. First, the computer 2 of the image recognition system 1 generates still image data 200 by combining two or more images 100 taken of an object at different times (step S1). Specifically, the computer 2 acquires multiple images 100 as still images of an object (for example, a moving person, car, or landscape) that has been photographed multiple times at different times, and generates a single still image data 200 by combining the acquired multiple still images. The two or more images 100 may be extracted from a video of the object at different time intervals or predetermined intervals.

[0041] Next, the computer 2 assigns a number to each of the two or more images 100 in order to indicate the chronological order in which the images were taken (step S2). Specifically, the computer 2 assigns numbers to the two or more images 100 included in the single still image data 200 generated in step S1 above, based on the order in which they were taken, for example, "1" for the first image taken, "2" for the next image taken, "3" for the next, and so on.

[0042] Next, computer 2 uses the still image data 200 and a predetermined prompt 300 to ask the generating AI a predetermined question about the situation and obtains an answer from the generating AI (step S3). Specifically, computer 2 combines the single still image data 200 generated in step S2 described above with a predetermined prompt (instruction) 300 for asking the generating AI, which is incorporated as part of the multimodal AI, a question about the situation of the image, and asks the generating AI a question. The generating AI then analyzes the contents of the still image data 200 and obtains an answer generated based on the prompt.

[0043] The "multimodal AI" mentioned above generally refers to AI technology that comprehensively understands and processes multiple modals (data formats: not limited to text characters, but including images, videos, programs, etc.). Furthermore, the "generative AI" mentioned above generally refers to generative artificial intelligence technology that generates new data or content (not limited to text characters, but including images, videos, programs, etc.) based on given input data, and can be implemented using a chatbot that utilizes a large-scale language model. "Generative AI" can generate new data not included in the training data by learning the regularities and structures of the training data during training. Therefore, the "Generative AI incorporated as part of multimodal AI" mentioned above refers to a technology that integrally processes different data formats and generates new data and content, and in this specification, it is simply referred to as "Generative AI".

[0044] The above is an overview of the processes performed by the image recognition system 1.

[0045] [System configuration of image recognition system 1] Based on Figure 2, the system configuration of the image recognition system 1 of this embodiment will be described. The image recognition system 1 consists of a computer 2 and is a computer system for recognizing the state of an object that changes over time.

[0046] The image recognition system 1 may include other terminals and devices. For example, each user may use a different computer 2, in which case the image recognition system 1 will perform the processes described later using one or more combinations of computer 2 and the other included terminals and devices.

[0047] Computer 2 may be implemented as, for example, a single terminal device, multiple terminal devices, or a virtual device like a cloud computer. Alternatively, it may be implemented in an on-premises environment using a dedicated server and network.

[0048] Computer 2 includes, for example, computers such as desktop PCs, laptops, and servers; mobile devices such as smartphones and tablet devices; and wearable devices such as head-mounted displays like smart glasses and smartwatches.

[0049] Computer 2 includes a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), RAM (Random Access Memory), ROM (Read Only Memory), and other components as its control unit.

[0050] Computer 2 includes data storage as a memory unit, such as a hard disk, semiconductor memory, recording medium, or memory card. The data storage may be internal data storage and / or external data storage. The data may be stored in a cloud service or database.

[0051] Computer 2 includes a communication unit equipped with devices to enable communication with other terminals and equipment. The communication method may be wireless or wired.

[0052] Computer 2 shall have the necessary functions for operating Computer 2 as an input unit. Examples of input methods include a liquid crystal display for touch panel functionality, a keyboard, a mouse, a pen tablet, hardware buttons on the device, a microphone for voice recognition, etc. The present invention is not particularly limited by the input method.

[0053] Computer 2 shall have the necessary functions for a user of the image recognition system 1 to operate Computer 2 as an output unit. Examples of output methods include display on an LCD display, a PC display, projection to a projector, and audio output. The present invention is not particularly limited by the output method.

[0054] Computer 2, as an imaging unit, is equipped with the necessary functions to capture images such as video and / or still images, but it may also acquire images captured by an external imaging device via the communication unit described above.

[0055] The control unit works in cooperation with the processing unit to realize the image generation unit 21 and the numbering unit 22. The control unit also works in cooperation with the processing unit and the communication unit to realize the questioning unit 23.

[0056] The above describes the system configuration of image recognition system 1.

[0057] [Image Recognition Processing] Based on Figure 3, the image recognition process performed by computer 2 will be described. Figure 3 is a flowchart of the image recognition process performed by computer 2. The image recognition process consists of steps S11 to S13, as shown in Figure 3, and is the process performed in steps S1 to S3 described above.

[0058] First, the image generation unit 21 of the computer 2 of the image recognition system 1 generates still image data 200 by combining two or more images 100 of the object taken at different times (step S11). Specifically, the image generation unit 21 acquires images of the object (for example, a moving person, car, or landscape) taken multiple times at different times as multiple images 100, and generates a single still image data 200 by combining the acquired multiple images 100. The two or more images may be extracted from a video of the object at different time intervals or predetermined intervals.

[0059] Figure 4 shows an example of 100 images of the same object, taken at different times. Figure 4 shows three images taken in the order of A, B, and C. In image A, ball YL is about to be placed into the box, which is the object. In image B, ball OR is about to be placed into the box, which is the object. In image C, ball BL is about to be placed into the box, which is the object.

[0060] Figure 5 shows an example of a single still image data 200 formed by combining the images shown in Figure 4. Figure 5 shows a single still image data 200 formed by combining the three images A, B, and C shown in Figure 4. As shown in Figure 5, the image generation unit 21 generates a still image data 200 by combining multiple images A, B, and C into a single image data. In Figure 5, images A, B, and C are combined vertically to form a single still image data 200, but any method of combining them is acceptable as long as it results in a single still image data 200.

[0061] In step S11, the image generation unit 21 of the computer 2 may perform noise reduction, resolution adjustment, or image normalization on two or more images 100 before generating the still image data 200. Specifically, before generating the still image data 200, the image generation unit 21 may, for example, remove unwanted noise from each image, adjust the resolution of each image to an appropriate size, or perform normalization to equalize image characteristics such as brightness and contrast on two or more images 100.

[0062] Furthermore, in step S11, if audio data, sensor data, or text data is associated with each of the two or more images 100, the image generation unit 21 of the computer 2 may integrate the audio data, sensor data, or text data into the still image data. Specifically, if audio data, sensor data, or text data is associated with each of the two or more images 100, the image generation unit 21 may integrate the associated audio data, sensor data, or text data into the two or more images 100 that are combined into the still image data 200.

[0063] Next, the numbering unit 22 of computer 2 assigns a number to each of the two or more images 100 in order to indicate the chronological order in which the images were taken (step S12). Specifically, the numbering unit 22 assigns numbers to the two or more images 100 that were combined with the single still image data 200 generated by the image generation unit 21 in step S11, based on the order in which they were taken, for example, "1" for the first image taken, "2" for the next image taken, "3" for the next, and so on. A timestamp may also be added to indicate the chronological order.

[0064] Figure 6 shows an example in which a temporal order is assigned to each image in the still image data 200 shown in Figure 5. As shown in Figure 6, the numbering unit 22 assigns the numbers "1" to image A, "2" to image B, and "3" to image C to the still image data 200, as shown in Figure 5, which was generated by the image generation unit 21 in step S11 described above. In Figure 6, the numbers are assigned to the lower right of each image A, B, and C, but the numbers can be assigned to any position.

[0065] Next, the questioning unit 23 of computer 2 uses the still image data 200 and a predetermined prompt 300 to ask the generating AI a predetermined question about the situation (step S13). Specifically, the questioning unit 23 combines the single still image data 200 generated in step S12 described above with a predetermined prompt (instruction) 300 for asking the generating AI a question about the situation of the image, asks the generating AI a question, and the generating AI analyzes the contents of the still image data 200 and obtains an answer generated based on the prompt.

[0066] For example, if a predetermined prompt 300 is set for a single still image data 200 as shown in Figure 6, and the phrase is, for example, "Please check if the box contains balls BA, BB, and BC in the order of balls YL, OR, BL," the questioning unit 23 provides this prompt to the generating AI along with the still image data 200. The generating AI then infers the situation that "the box contains balls BA, BB, and BC in the order of balls YL, OR, BL" from the temporally consecutive images A, B, and C of the still image data 200, and generates an answer. The prompt 300 may be stored as data in the memory of the computer 2 in advance, or it may be set as data by the user via the input unit of the computer 2.

[0067] The specified prompt 300 may appropriately include a "role" for answering, specified "constraints," and an "output format." The "role" specifies the position or perspective from which the generating AI should answer. This ensures that the tone and content of the answer are appropriate to the role the generating AI takes on. The "constraints" specify the rules and conditions that the generating AI must follow when creating an answer. This limits the scope, style, and content of the answer, resulting in output that aligns with the user's intent. The "output format" specifies the format in which the generating AI should output the answer. The "constraints" may also include answers regarding the previous situation. This allows the generating AI to understand the current situation by referring to the situation answered immediately before, and to generate an answer that demonstrates temporal continuity from the past to the present.

[0068] Figure 7 shows another example of a single still image data set 200 in which each image is assigned a temporal order. In Figure 7, for example, in the image numbered "13", there are three stuffed animals YL, BL, and RD inside a box; in the image numbered "14", stuffed animal GR is being held up by a person's hand on top of the box containing the three stuffed animals YL, BL, and RD; in the image numbered "15", stuffed animal GR is released from the person's hand and placed inside the box containing the three stuffed animals YL, BL, and RD; and in the image numbered "16", stuffed animal GR is inside the box together with the three stuffed animals YL, BL, and RD.

[0069] For example, if the "Role" to be included in a predetermined prompt 300 for such still image data 200 is set to "You are an AI assistant that explains the content of a video. Infer and explain the situation in the given video frame," and the "Constraints" are set to include the answer regarding the previous situation, such as "Each input image is divided into frames in a tile-like pattern with a black border. Read the numbers in the upper left corner of the frames in order. Infer the situation from the previous frame. However, give priority to the facts that can be read from the frame. For example, if a stuffed animal is about to be put into a box but the situation of it being put in cannot be confirmed, interpret that it has not been put in," and the "Output Format" is set to "The output should be in the format of "Number read from the upper left corner of the frame: Description of the frame" for each frame," then according to these "Role," "Constraints," and "Output Format," the generating AI will limit the range, style, and content of its answers and produce output that aligns with the user's intent.

[0070] Furthermore, in this example of prompt 300, the "limitations" include the statement, "Please infer from the situation in the previous frame. However, give priority to facts that can be read from the frame. For example, if it is not possible to confirm that a stuffed animal has been placed in a box, even if it appears to be about to be placed in the box, interpret that it has not been placed in the box." As a result, the generating AI infers a temporally continuous situation from the past, "stuffed animals YL, BL, and RD were in the box," to the present, "stuffed animal GR was placed in the box containing stuffed animals YL, BL, and RD," and generates an answer based on this.

[0071] In step S13, the questioning unit 23 of computer 2 may ask multiple generating AIs predetermined questions about the situation and integrate the answers obtained from these generating AIs. Specifically, the image generation unit 21 may ask multiple generating AIs predetermined questions about the situation individually, collect the answers obtained from each generating AI, integrate those answers, and derive a single unified answer as a whole.

[0072] Furthermore, in step S13, the questioning unit 23 of computer 2 may output still image data 200 and the answer from the generating AI, and may receive predetermined questions and feedback from the user. In this case, in step S11, the image generation unit 21 of computer 2 may select two or more images 100 based on the answer or feedback received in step S13 executed in the previous process, and generate still image data 200. Specifically, in step S13, the questioning unit 23 may output the answer obtained from the still image data 200 and the generating AI to the output unit of computer 2, and may receive predetermined questions and feedback from the user via the input unit of computer 2. In this case, in step S11, the image generation unit 21 of computer 2 may select an appropriate image from two or more images 100 based on the content of the answer or feedback received from the user in step S13 executed in the previous process, and generate still image data 200 by combining the selected images.

[0073] The still image data 200 and the response of the generating AI output to the output unit of computer 2 in step S13 described above may be managed in either an on-premise computing environment or a cloud computing environment, depending on the user's preference.

[0074] The above describes the image recognition process.

[0075] Therefore, according to the image recognition system 1, by combining multiple images taken at different times and assigning numbers to the captured images to indicate their temporal order, it becomes possible to provide a foundation for the generative AI to accurately understand the situation. By asking the generative AI questions using these multiple images and predetermined prompts, the AI ​​can detect changes and patterns that humans tend to overlook.

[0076] Furthermore, according to image recognition system 1, by including a role in the prompt given to the generating AI, the position and perspective from which the generating AI responds becomes clear, the direction of the response is determined, and ambiguous or inappropriate content can be prevented. In addition, by including constraints in the prompt, unnecessary information or misleading content can be eliminated. Moreover, by specifying the output format in the prompt, consistency in the response is maintained, and information can be obtained in a format that is easy for the user to use directly.

[0077] Furthermore, according to the image recognition system 1, by using still images extracted at predetermined intervals as the still images provided to the generating AI, the generating AI can more easily analyze temporal changes and track them more readily.

[0078] Furthermore, according to image recognition system 1, by generating new information based on the inference obtained as the answer immediately beforehand, the analysis flow is not interrupted, and consistent results can be obtained, making it possible to obtain a consistent answer.

[0079] Furthermore, according to the image recognition system 1, removing unwanted noise from images improves visual quality and increases the accuracy of generating still image data. Adjusting the resolution also allows for optimal size optimization of the generated still image data while preserving image detail. Additionally, image normalization unifies the brightness and contrast of images, enabling more accurate recognition of differences between images.

[0080] Furthermore, according to image recognition system 1, by integrating responses from multiple generative AIs, it becomes possible to compensate for the weaknesses of individual AIs and obtain more accurate and reliable responses.

[0081] Furthermore, according to the image recognition system 1, the addition of audio data and text data such as descriptive text related to the image clarifies the intent and content of the image. In addition, the provision of environmental information (temperature, humidity, location information, etc.) from sensor data related to the image provides a more detailed and clearer understanding of the background and circumstances of the image. As a result, more accurate judgments and predictions can be made by the generating AI.

[0082] Furthermore, according to the image recognition system 1, by outputting still image data and the generating AI's response while accepting questions and feedback from the user, the user can not only obtain the necessary information but also reflect their own opinions and questions in the questions and feedback to the generating AI.

[0083] Furthermore, according to the image recognition system 1, by outputting still image data and the generating AI's response while accepting questions and feedback from the user, the user can not only obtain the necessary information but also reflect their own opinions and questions in the questions and feedback to the generating AI.

[0084] Furthermore, according to the image recognition system 1, it becomes possible to manage data by selecting an environment that suits the user's needs, such as selecting an on-premises computing environment when security requirements are strict, a cloud computing environment when scalability is required, or performing specific processing in an on-premises computer environment and other processing in a cloud computing environment.

[0085] The means and functions described above are realized by a computer (including the CPU, information processing unit, and various terminals) reading and executing a predetermined program. The program is provided, for example, from one or more computers via a network (cloud service, SaaS: Software as a Service). Alternatively, the program may be provided, for example, recorded on a computer-readable recording medium. In this case, the computer reads the program from the recording medium, transfers it to an internal or external recording device, records it, and executes it. Alternatively, the program may be pre-recorded on a recording device (recording medium) such as a magnetic disk, optical disk, or magneto-optical disk, and provided to the computer from that recording device via a communication line.

[0086] Although embodiments of the present invention have been described above, the present invention is not limited to these embodiments. Furthermore, the effects described in the embodiments of the present invention are merely a list of the most preferred effects arising from the present invention, and the effects of the present invention are not limited to those described in the embodiments. [Explanation of symbols]

[0087] 1 Image recognition system, 2 Computer, 21 Image generation unit, 22 Numbering unit, 23 Question unit, 100 Images, 200 Still image data, 300 Prompts

Claims

1. An image recognition system that recognizes the state of an object that changes over time, An image recognition system comprising a questioning unit that uses still image data, which is a single image data obtained by combining two or more images of the aforementioned object taken at different times, and which is numbered to indicate the temporal order in which the two or more images included in the still image data were taken, and a predetermined prompt, to ask a predetermined question to a generating AI regarding the aforementioned situation and to obtain an answer from the generating AI.

2. The aforementioned predetermined prompt includes a role for responding, predetermined constraints, and an output format. The image recognition system according to claim 1.

3. The aforementioned still image data is extracted at predetermined intervals from a video of the subject object. The image recognition system according to claim 1.

4. The aforementioned predetermined constraints further include the aforementioned answers obtained, The image recognition system according to claim 2.

5. Before the aforementioned still image data is generated, noise reduction, resolution adjustment, or image normalization is performed on the two or more images. The image recognition system according to claim 1.

6. The questioning unit asks a plurality of generating AIs the predetermined questions regarding the situation, obtains answers from the plurality of generating AIs and integrates them. The image recognition system according to claim 1.

7. If audio data, sensor data, or text data is associated with each of the two or more images, the audio data, sensor data, or text data is integrated into the still image data. The image recognition system according to claim 1.

8. The questioning unit outputs the still image data and the answer generated by the AI, and receives the predetermined questions and feedback from the user. The image recognition system according to claim 1.

9. Based on the received response or feedback, the selection of two or more images is made, and the still image data is generated. The image recognition system according to claim 8.

10. The system runs in either an on-premises computing environment or a cloud computing environment, and the static image data and the AI-generated responses are managed in an environment desired by the user. The image recognition system according to any one of claims 1 to 9.

11. An image recognition method performed by a computer to recognize the state of an object that changes over time, The step of using a still image data, which is a still image data obtained by combining two or more images of the aforementioned object taken at different times into a single image data, and in order to indicate the temporal order in which the two or more images included in the still image data are taken, a number is assigned to each of the two or more images, and a predetermined prompt, to ask a predetermined question to the generating AI regarding the aforementioned situation and to obtain a response from the generating AI, An image recognition method comprising the following features.

12. On the computer, A still image data in which two or more images of an object are taken at different times and combined into a single image data, and in order to indicate the chronological order in which the two or more images included in the still image data are taken, a number is assigned to each of the two or more images in the still image data, and a predetermined prompt is used to ask a predetermined question about the situation to a generating AI and to obtain a response from the generating AI. A computer-readable program for executing a command.