Image recognition system, method, and program
The image recognition system addresses the limitations of generative AI by combining frame images with timestamps and prompts, enhancing temporal understanding and accuracy in recognizing changing objects.
Patent Information
- Application Number
- JP2025134881
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Generative AI systems struggle to recognize objects that change over time due to their inability to understand temporal relationships and process large volumes of continuous image data effectively, leading to difficulties in advanced situation recognition.
An image recognition system that combines multiple frame images from different times into a single image, assigns timestamps, and uses prompts to ask generation AI questions, incorporating audio, sensor, and text data to enhance understanding and accuracy.
Enables generation AI to accurately recognize changes and patterns over time, providing consistent and reliable answers by integrating multiple AI responses and managing data in user-defined environments.
Smart Images

Figure 0007767682000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image recognition system, an image recognition method, and a program for recognizing an object that changes over time. [Background technology]
[0002] Recent image recognition technologies incorporate generative AI as part of multimodal AI to identify specific objects or patterns in images. Specifically, multimodal AI, such as large-scale language models (LLMs), which have the ability to integrate and process multiple data formats such as text, images, and audio, complements generative AI, which generates new data based on input data, analyzing images and generating explanatory text.
[0003] For example, Non-Patent Document 1 generates text output (natural language, code, etc.) for inputs that contain a mixture of text and images. Also, Non-Patent Document 2, for example, provides information that deepens a user's understanding of what they see in their pass-through environment. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] GPT-4, [online], [Retrieved December 16, 2024], Internet,<https: / / openai.com / index / gpt-4-research / > [Non-patent document 2] About Meta AI in Meta Quest, [online], [Retrieved December 16, 2024], Internet, <https: / / www.meta.com / ja-jp / help / quest / articles / in-vr-experiences / oculus-features / learn-mBe-meta-ai-meta-quest / > Summary of the Invention [Problem to be solved by the invention]
[0005] However, generative AI processes input data as individual, static pieces of information, and is therefore unable to directly understand changes over time or causal relationships. For example, even if a series of images is given to it, it cannot recognize them as a "time flow," and each image is treated as an independent piece of information.
[0006] In addition, there is a limit to the amount of information that generative AI can process. For example, in manufacturing sites where each worker's work status is recorded on video using one or more cameras, it is difficult to import the huge amount of video data corresponding to the number of cameras all at once, and the data must be thinned out or summarized as necessary.
[0007] Due to these constraints, when generative AI is asked to process continuous image data, it is generally possible to recognize the objects and features contained in each image (image recognition) by dividing the data into time slices (still images separated by time) and loading each one individually, but because it has no concept of time, it is unable to understand the continuity or causal relationships between images.
[0008] For example, even if it is recognized in one image that "a worker has put something into a box," it is extremely difficult to link this with "an object being moved" in the next image and grasp the overall situation (situation recognition), such as "the object is moving as a result of being put in." Therefore, in order to achieve advanced situation recognition that takes the time axis into consideration, it is necessary to combine generative AI with a system specialized in temporal data processing, rather than using it alone.
[0009] In view of these problems, the present invention aims to provide an image recognition system, an image recognition method, and a program that allows a generating AI to recognize objects that change over time in two or more frame images by assigning times to two or more frame images extracted at different times from each video shot from one or more directions for a predetermined period of time, combining them into a single image, and entering the chronological order in a prompt. [Means for solving the problem]
[0010] The present invention provides the following solutions.
[0011] According to the first aspect of the invention, An image recognition system that recognizes the situation of an object that changes over time, an image generating unit that generates still image data by combining two or more frame images extracted at different times from each video image obtained by photographing the object from one or more directions for a predetermined period of time; a time assigning unit that assigns a time to each of the two or more frame images included in the generated still image data in order to indicate the chronological order in which the two or more frame images were captured; A questioning unit that uses the still image data and a predetermined prompt to ask the generation AI a predetermined question about the situation and obtains an answer from the generation AI; An image recognition system is provided.
[0012] According to the invention relating to the first feature, by combining multiple frame images extracted at different times from each video shot from one or more directions for a predetermined period of time and assigning a time stamp to the extracted frame images, it is possible to provide a foundation for the generation AI to accurately understand the situation, and by asking the generation AI questions using these multiple frame images and predetermined prompts, it becomes possible for the AI to detect changes and patterns that humans tend to overlook.
[0013] The second aspect of the invention is the first aspect of the invention, the predetermined prompt includes a role for answering, predetermined constraints, and an output format; Provides an image recognition system.
[0014] According to the second aspect of the invention, by including a role in the prompt given to the generation AI, the position and perspective from which the generation AI responds become clear, the direction of the response is determined, and it becomes possible to prevent ambiguous responses or inappropriate content. Furthermore, by including constraints in the prompt, it becomes possible to eliminate unnecessary information or content that may lead to misunderstandings. Furthermore, by specifying the output format in the prompt, it becomes possible to maintain consistency in the response and obtain information in a format that is easy for the user to use as is.
[0015] The third aspect of the invention is the first aspect of the invention, The still image data is extracted at predetermined intervals from a video of the object. Provides an image recognition system.
[0016] According to the third feature of the invention, by using still images extracted at predetermined intervals as still images to be given to the generation AI, it becomes easier for the generation AI to analyze changes over time and to track them easily.
[0017] The fourth aspect of the invention is the second aspect of the invention, The predetermined constraints further include the obtained answer. Provides an image recognition system.
[0018] According to the invention relating to the fourth feature, by generating new information based on the inference of the answer obtained immediately before, the flow of analysis is not interrupted and consistent results are obtained, making it possible to obtain consistent answers.
[0019] The fifth aspect of the invention is the first aspect of the invention, the image generation unit performs noise removal, resolution adjustment, or image normalization on the two or more frame images before generating the still image data; Provides an image recognition system.
[0020] According to the fifth aspect of the invention, by removing unnecessary noise contained in frame images, visual quality is improved and the accuracy of generating still image data is increased. Furthermore, by adjusting the resolution, it is possible to optimize the size of the generated still image data while maintaining image detail. Furthermore, by normalizing the frame images, the brightness and contrast of the frame images are unified, making it possible to more accurately recognize differences between images.
[0021] The sixth aspect of the invention is the first aspect of the invention, The questioning unit asks the predetermined questions regarding the situation to a plurality of generation AIs, and obtains and integrates answers from the plurality of generation AIs. Provides an image recognition system.
[0022] According to the sixth feature of the invention, by integrating answers from multiple generation AIs, it is possible to compensate for the weaknesses of individual AIs and obtain more accurate and reliable answers.
[0023] The seventh aspect of the invention is the first aspect of the invention, When audio data, sensor data, or text data is associated with each of the two or more frame images, the image generation unit integrates the audio data, the sensor data, or the text data into the still image data. Provides an image recognition system.
[0024] According to the seventh feature of the invention, the intention and content of the frame image are clarified by adding text data such as audio data and explanatory text related to the frame image. Furthermore, environmental information (temperature, humidity, location information, etc.) from sensor data related to the frame image is provided, making the background and situation of the frame image clearer in more detail. This allows the generation AI to make more accurate judgments and predictions.
[0025] The eighth aspect of the invention is the first aspect of the invention, The questioning unit outputs the still image data and the answer of the generation AI, and receives the predetermined question and feedback from the user. Provides an image recognition system.
[0026] According to the eighth feature of the invention, by outputting still image data and the generation AI's answers while accepting questions and feedback from the user, the user not only obtains the information they need, but also has the ability to reflect their own opinions and doubts in the questions and feedback to the generation AI.
[0027] The ninth aspect of the invention is the eighth aspect of the invention, the image generation unit selects the two or more frame images based on the received answer or feedback, and generates the still image data. Provides an image recognition system.
[0028] According to the invention relating to the ninth feature, by outputting still image data and the generation AI's answers while accepting questions and feedback from the user, the user not only obtains the information they need, but also has the ability to reflect their own opinions and doubts in the questions and feedback to the generation AI.
[0029] The invention according to a tenth feature is the invention according to any one of the first to ninth features, The system is executed in either an on-premise computing environment or a cloud computing environment, and the still image data and the answers of the generating AI are managed in the user's desired environment. Provides an image recognition system.
[0030] According to the invention relating to the tenth feature, it becomes possible to manage data by selecting an environment that meets the user's needs, for example, by selecting an on-premise computing environment when security requirements are strict and a cloud computing environment when scalability is required, or by performing certain processes in an on-premise computer environment and other processes in a cloud computing environment.
[0031] Although the present invention is in the category of computer systems, similar actions and effects according to the category can also be achieved in other categories such as methods and programs. [Effects of the Invention]
[0032] According to the present invention, by extracting two or more frame images from videos taken from one or more directions for a predetermined period of time, assigning times to the extracted frame images at different times, combining them into a single image, and entering the chronological order in a prompt, it is possible to provide an image recognition system, image recognition method, and program that allows a generating AI to recognize objects that change over time in the two or more frame images. [Brief explanation of the drawings]
[0033] [Figure 1] FIG. 1 is a diagram illustrating an overview of an image recognition system 1 according to an embodiment of the present invention. [Figure 2] 1 is a configuration diagram of an image recognition system 1 according to an embodiment of the present invention. [Figure 3] 1 is a flowchart of an image recognition process executed by the image recognition system 1 of the present embodiment. [Figure 4] 1A and 1B are diagrams showing how an object is photographed from various directions in the image recognition system 1 of this embodiment. [Figure 5] FIG. 1 is a diagram showing an example of a frame image 100 of a moving image in which an object is photographed from each direction in the image recognition system 1 of this embodiment. [Figure 6] FIG. 2 is a diagram showing an example of one piece of still image data 200 formed by combining frame images 100 in the image recognition system 1 of this embodiment. [Figure 7] FIG. 10 is a diagram showing an example in which time is added to each frame image 100 in still image data 200 in the image recognition system 1 of this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0034] The best mode for carrying out the present invention will be described below with reference to the drawings. However, this is merely an example, and the technical scope of the present invention is not limited to this example.
[0035] [Image Recognition System 1 Overview] An overview of an image recognition system 1 according to one embodiment of the present invention will be described with reference to Fig. 1. Fig. 1 is a diagram for explaining the overview of the image recognition system 1 according to one embodiment of the present invention. The image recognition system 1 is made up of a computer 2, and is a computer system for recognizing the situation of an object that changes over time.
[0036] The computer 2 of the image recognition system 1 is, for example, a computer such as a desktop computer, a laptop computer, or a server, a mobile terminal such as a smartphone or a tablet terminal, or a wearable terminal such as a head-mounted display such as smart glasses or a smart watch.
[0037] The computer 2 of the image recognition system 1 may be realized, for example, by one terminal device, by multiple terminal devices, or by a virtual device such as a cloud computer. Also, the computer 2 may be realized in an on-premise environment using a dedicated server or network.
[0038] Moreover, the image recognition system 1 may be configured with the above-mentioned terminal device instead of the computer 2.
[0039] The computer 2 of the image recognition system 1 is connected to the above-mentioned terminal device and other terminals and devices via a public line network or the like so as to be able to communicate data, and transmits and receives necessary data and information.
[0040] Next, an overview of the processing executed by the image recognition system 1 will be described. First, the computer 2 of the image recognition system 1 generates a single piece of image data 200 by combining two or more frame images 100 extracted at different times from each video captured from one or more directions for a predetermined period of time (step S1). Specifically, the computer 2 acquires, as still images, a plurality of frame images 100 extracted at different times from each video captured by one or more cameras of the working status of a worker at a manufacturing site, for example, and combines the acquired plurality of still images to generate a single piece of still image data 200. The subject being captured is the working status of a worker, but is not particularly limited to this.
[0041] Next, the computer 2 assigns a time to each of the two or more frame images 100 in order to indicate the chronological order in which the two or more frame images 100 included in the generated still image data 200 were captured (step S2). Specifically, the computer 2 assigns a time calculated based on the index of each frame image and the frame rate to the two or more frame images 100 included in the single still image data 200 generated in step S1 above.
[0042] Next, the computer 2 uses the still image data 200 and a predetermined prompt 300 to ask the generation AI a predetermined question about the situation and obtains an answer from the generation AI (step S3). Specifically, the computer 2 combines the single piece of still image data 200 generated in the above-mentioned step S2 with a predetermined prompt (instruction sentence) 300 for asking the generation AI, which is incorporated as part of the multimodal AI, a question about the situation of the image, and asks the generation AI, and the generation AI analyzes the contents of the still image data 200 and obtains an answer generated based on the prompt.
[0043] The aforementioned "multimodal AI" generally refers to AI technology that comprehensively understands and processes multiple modalities (data formats: not limited to text characters, but including images, videos, programs, etc.). The aforementioned "generative AI" generally refers to generative artificial intelligence technology that generates new data or content (not limited to text characters, but including images, videos, programs, etc.) based on given input data, and may be realized in chatbots that utilize large-scale language models. "Generative AI" can generate new data that is not included in the training data by machine learning the regularities and structures of the training data during training. Therefore, the above-mentioned "generative AI incorporated as part of multimodal AI" refers to a technology that processes different data formats in an integrated manner to generate new data and content, and is referred to simply as "generative AI" in this specification.
[0044] The above is an overview of the processing executed by the image recognition system 1.
[0045] [System configuration of image recognition system 1] The system configuration of the image recognition system 1 of this embodiment will be described with reference to Fig. 2. The image recognition system 1 is made up of a computer 2, and is a computer system for recognizing the situation of an object that changes over time.
[0046] It should be noted that the image recognition system 1 may include other terminals, devices, etc. For example, a different computer 2 may be used for each user, in which case the image recognition system 1 will execute each process described below using the computer 2 and / or a combination of the other included terminals, devices, etc.
[0047] The computer 2 may be realized, for example, by one terminal device, by multiple terminal devices, by a virtual device such as a cloud computer, or by a dedicated server or network in an on-premise environment.
[0048] The computer 2 may be, for example, a computer such as a desktop personal computer, a laptop computer, or a server, a mobile terminal such as a smartphone or a tablet terminal, or a wearable terminal such as a head-mounted display such as smart glasses or a smart watch.
[0049] The computer 2 includes a control unit such as a central processing unit (CPU), a graphics processing unit (GPU), a random access memory (RAM), and a read only memory (ROM).
[0050] The computer 2 includes a data storage unit such as a hard disk, semiconductor memory, recording medium, or memory card. The data storage may be internal data storage and / or external data storage. The data may be stored in a cloud service, a database, or the like.
[0051] The computer 2 includes a communication unit that is a device for enabling communication with other terminals, devices, etc. The communication method may be wireless or wired.
[0052] The computer 2 is assumed to have, as an input unit, functions necessary for operating the computer 2. Examples of devices for realizing input include an LCD display that realizes a touch panel function, a keyboard, a mouse, a pen tablet, hardware buttons on the device, and a microphone for voice recognition. The present invention is not particularly limited in function depending on the input method.
[0053] The computer 2 is an output unit that has the functions necessary for the user of the image recognition system 1 to operate the computer 2. Examples of output methods include display on an LCD display, a PC display, projection on a projector, and audio output. The present invention is not particularly limited in function by the output method.
[0054] The computer 2 has the necessary functions to capture images such as moving images and / or still images as an imaging unit, but may also acquire images captured by an external imaging device via the communication unit described above.
[0055] The control unit cooperates with the processing unit to realize an image generation unit 21 and a time setting unit 22. The control unit also cooperates with the processing unit and the communication unit to realize an interrogation unit 23.
[0056] The system configuration of the image recognition system 1 has been described above.
[0057] [Image recognition processing] The image recognition processing executed by the computer 2 will be described with reference to Fig. 3. Fig. 3 is a diagram showing a flowchart of the image recognition processing executed by the computer 2. As shown in Fig. 3, the image recognition processing is made up of steps S11 to S13, and is the processing executed in the above-mentioned steps S1 to S3.
[0058] First, the image generation unit 21 of the computer 2 of the image recognition system 1 generates still image data 200 by combining two or more frame images 100 extracted at different times from each video captured from one or more directions for a predetermined period of time to generate a single image data (step S11). Specifically, the image generation unit 21 reads each frame image 100 from each video captured by one or more cameras for a predetermined period of time, showing the working status of a worker at a manufacturing site or the like, assigns an index to each frame based on the order in which they were read, extracts multiple frame images 100 at different times based on the assigned index, acquires them as still images, and combines the acquired multiple still images to generate a single image data 200. In this example, the shooting location is a manufacturing site, and the shooting subject is the working status of a worker, but this is not particularly limited.
[0059] FIG. 4 is a diagram showing how an object is photographed from various directions. FIG. 4 shows how a worker is putting items into a box, being photographed by three cameras (camera A, camera B, and camera C). These cameras are shooting video of the worker putting parts A and B into the box in that order for a predetermined period of time. Camera A is shooting video of the worker's work from above. Camera B is shooting video of the worker's work from the front. Camera C is shooting video of the worker's work from the side.
[0060] Fig. 5 is a diagram showing examples of frame images 100 from a video in which an object is captured from each direction. Fig. 5 shows frame images 100 from each video in which a worker is putting items into a box and the situation is captured for 9 seconds by cameras (camera A, camera B, and camera C) installed in each direction at a frame rate of 30 fps. The frame images 100 shown in Fig. 5 are assigned index numbers #45, #90, #135, #180, #225, and #270, but in reality, each video is made up of frame images 100 assigned index numbers #1 to #270.
[0061] FIG. 6 is a diagram showing an example of a single piece of still image data 200 in which frame images 100 shown in FIG. 5 are combined, each extracted three seconds apart. FIG. 6 shows a single piece of still image data 200 in which nine frame images 100 with indexes #90, #180, and #270 extracted three seconds apart from the video captured by each camera (camera A, camera B, and camera C) shown in FIG. 5 are combined together. As shown in FIG. 6, the image generation unit 21 generates a single piece of still image data 200 by combining, for example, a plurality of frame images 100 extracted three seconds apart from the frame images 100 of each video. In FIG. 6, the frame images 100 for each camera (camera A, camera B, and camera C) are combined vertically in index order to form a single piece of still image data 200, but they may be combined in any way to form a single piece of still image data 200.
[0062] In step S11, the image generation unit 21 of the computer 2 may perform noise removal, resolution adjustment, or image normalization on two or more frame images 100 before generating the still image data 200. Specifically, before generating the still image data 200, the image generation unit 21 may perform normalization on two or more frame images 100, for example, to remove unnecessary noise in each image, adjust the resolution of each image to an appropriate size, or normalize image characteristics such as brightness and contrast.
[0063] Furthermore, in step S11, the image generation unit 21 of the computer 2 may integrate the audio data, sensor data, or text data into the still image data if audio data, sensor data, or text data is associated with each of the two or more frame images 100. Specifically, if audio data, sensor data, or text data is associated with each of the two or more frame images 100, the image generation unit 21 may integrate the associated audio data, sensor data, or text data into the two or more frame images 100 to be combined into the still image data 200.
[0064] Next, the time assignment unit 22 of the computer 2 assigns a time to each of the two or more frame images 100 included in the generated still image data 200, in order to indicate the chronological order in which the two or more frame images 100 were captured (step S12). Specifically, for the two or more frame images 100 combined into one still image data 200 generated by the image generation unit 21 in step S11 described above, if the frame rate of the captured video is 30 fps and the index of the extracted frame image is #180, the time assignment unit 22 assigns the time of this frame image to 180 / 60=3 (seconds). This is performed for each frame image. The method of assigning a time to each frame image is not limited to this, and time codes (PTS (Presentation Time Stamp), DTS (Decoding Time Stamp)) included in the metadata may also be used.
[0065] Fig. 7 is a diagram showing an example in which a time is assigned to each image in the still image data 200 shown in Fig. 6. As shown in Fig. 7, the time assignment unit 22 calculates the time of each frame image 100 based on the index of each frame image 100 and the frame rate for still image data 200 such as that shown in Fig. 6 generated by the image generation unit 21 in step S11 described above, and assigns the time "00:00:03" to the frame image 100 with index #90, "00:00:06" to the frame image 100 with index #180, and "00:00:09" to the frame image 100 with index #270. In Fig. 7, the time is assigned below each of images A, B, and C, but the time may be assigned anywhere.
[0066] Next, the interrogation unit 23 of the computer 2 asks the generation AI a predetermined question about the situation using the still image data 200 and a predetermined prompt 300 (step S13). Specifically, the interrogation unit 23 combines the single piece of still image data 200 generated in the above-mentioned step S12 with a predetermined prompt (instruction sentence) 300 for asking the generation AI a question about the situation of the image, and asks the generation AI a question, and the generation AI analyzes the contents of the still image data 200 and obtains an answer generated based on the prompt.
[0067] For example, if a predetermined prompt 300 is set for a single piece of still image data 200 as shown in Fig. 7, such as "Please confirm that parts A and B are placed in the box in this order," the questioning unit 23 provides this prompt along with the still image data 200 to the generating AI, causing the generating AI to infer the situation of "parts A and B are placed in the box in this order" from the frame image 100 of the still image data 200 and generate this as an answer. Note that the prompt 300 may be stored in advance as data in the storage unit of the computer 2, or may be set as data by the user via the input unit of the computer 2.
[0068] The predetermined prompt 300 may include a "role" for answering, predetermined "constraints," and an "output format" as appropriate. The "role" specifies the position or perspective from which the generation AI should answer. As a result, the generation AI has a specific role, and the tone and content of the answer will be appropriate for that role. The "constraints" specify the rules and conditions that the generation AI must follow when creating an answer. This limits the range, style, and content of the answer, resulting in an output that meets the user's intentions. The "output format" specifies the format in which the generation AI should output the answer. Note that the "constraints" may also include an answer related to the previous situation. This allows the generation AI to understand the current situation by referring to the situation in the previous answer and generate an answer that is an inference with temporal continuity from the past to the present.
[0069] In FIG. 7, in the three frame images 100 assigned the time "00:00:03," the worker has part A held up directly above the box in his right hand, and part B in his left hand. In the three frame images 100 assigned the time "00:00:06," part B is held up in the worker's left hand, at the upper right corner as viewed from the front of the box, and in the frame image 100 taken by camera A, part A is inside the box. In the three frame images 100 assigned the time "00:00:09," the worker's left hand is inside the box, and in the frame image 100 taken by camera A, the worker's left hand holding part B is inside the box containing part A.
[0070] For example, for such still image data 200, the "role" included in a predetermined prompt 300 is set to "You are an AI assistant that explains the contents of the video. Please infer and explain the situation in the given video frame." The "constraints" are set to include answers regarding the immediately preceding situation, with the following wording: "Each input image is divided into frames in a tiled pattern by a black border. Please read the frames in the order of the time at the bottom. Please infer from the situation in the previous frame. However, please give top priority to the facts that can be read from the frame. For example, if a part is about to be placed in a box but the situation in which it has been placed cannot be confirmed, please interpret that it has not been placed." The "output format" is set to "The output for each frame should be in the format: 'Time read from the bottom of the frame: Description of the frame'" The range, style, and content of the answers from the generating AI are limited according to these "role," "constraints," and "output format," and the output is in line with the user's intentions.
[0071] In this example of prompt 300, the "restrictions" include the following text: "Please make an inference from the situation in the previous frame. However, give top priority to the facts that can be read from the frame. For example, if a part appears to be about to be placed in a box but the situation in which it has been placed cannot be confirmed, then interpret it as not having been placed." As a result, in the still image data 200 shown in Figure 7, the generation AI infers a situation with time continuity from the past, in which "part A was placed in the box," to the present, in which "part B has not been placed in the box that contains part A," and generates an answer.
[0072] In step S13, the questioning unit 23 of the computer 2 may ask a plurality of generation AIs a predetermined question about the situation, obtain and integrate answers from the plurality of generation AIs. Specifically, the image generating unit 21 may individually ask a plurality of generation AIs a predetermined question about the situation, collect answers obtained from each generation AI, integrate those answers, and derive a single unified answer as a whole.
[0073] Furthermore, in step S13, the interrogation unit 23 of the computer 2 may output the still image data 200 and the answer of the generation AI, and may accept a predetermined question and feedback from the user. In this case, in step S11, the image generation unit 21 of the computer 2 may select two or more frame images 100 based on the answer or feedback received in step S13, which was executed in the previous process, and generate the still image data 200. Specifically, in step S13, the interrogation unit 23 may output the still image data 200 and the answer obtained from the generation AI to the output unit of the computer 2, and may accept a predetermined question or feedback from the user via the input unit of the computer 2. In this case, in step S11, the image generation unit 21 of the computer 2 may select appropriate images from two or more frame images 100 based on the content of the answer or feedback received from the user in step S13, which was executed in the previous process, and generate the still image data 200 by combining the selected images.
[0074] The still image data 200 and the response of the generated AI output to the output section of the computer 2 in step S13 described above may be managed in either an on-premise computing environment or a cloud computing environment, depending on the user's wishes.
[0075] This completes the image recognition process.
[0076] Therefore, according to image recognition system 1, by combining multiple frame images extracted at different times from each video shot from one or more directions for a predetermined period of time and assigning a time stamp to the extracted frame images, it is possible to provide a foundation for the generation AI to accurately understand the situation, and by asking the generation AI questions using these multiple frame images and predetermined prompts, the AI can detect changes and patterns that humans tend to overlook.
[0077] Furthermore, according to the image recognition system 1, by including a role in the prompt given to the generation AI, the position and perspective from which the generation AI responds become clear, the direction of the response is determined, and it becomes possible to prevent ambiguous responses or inappropriate content. Also, by including constraints in the prompt, it becomes possible to eliminate unnecessary information or content that may lead to misunderstandings. Furthermore, by specifying the output format in the prompt, it becomes possible to maintain consistency in the response and obtain information in a format that is easy for users to use as is.
[0078] Furthermore, according to the image recognition system 1, by using still images extracted at predetermined intervals as still images to be provided to the generation AI, it becomes easier for the generation AI to analyze changes over time and to track them easily.
[0079] Furthermore, according to the image recognition system 1, by generating new information based on the inference of the answer obtained immediately before, the flow of analysis is not interrupted and consistent results are obtained, making it possible to obtain consistent answers.
[0080] Furthermore, the image recognition system 1 improves visual quality and increases the accuracy of generating still image data by removing unnecessary noise contained in frame images. Also, by adjusting the resolution, it becomes possible to optimize the size of the generated still image data while preserving image detail. Furthermore, normalizing frame images makes it possible to unify the brightness and contrast of frame images, enabling more accurate recognition of differences between images.
[0081] Furthermore, according to the image recognition system 1, by integrating answers from multiple generation AIs, it is possible to compensate for the weaknesses of individual AIs and obtain more accurate and reliable answers.
[0082] Furthermore, with the image recognition system 1, the intention and content of the frame image are clarified by adding text data such as audio data and explanatory text related to the frame image. Furthermore, environmental information (temperature, humidity, location information, etc.) is provided from sensor data related to the frame image, making the background and situation of the frame image clearer in more detail. As a result, the generation AI can make more accurate judgments and predictions.
[0083] Furthermore, according to image recognition system 1, by outputting still image data and the generation AI's answers while accepting questions and feedback from the user, the user can not only obtain the information they need, but also reflect their own opinions and doubts in the questions and feedback to the generation AI.
[0084] Furthermore, according to image recognition system 1, by outputting still image data and the generation AI's answers while accepting questions and feedback from the user, the user can not only obtain the information they need, but also reflect their own opinions and doubts in the questions and feedback to the generation AI.
[0085] Furthermore, according to the image recognition system 1, it is possible to manage data by selecting an environment that meets the user's needs, for example, by selecting an on-premise computing environment when security requirements are strict and a cloud computing environment when scalability is required, or by performing certain processing in an on-premise computer environment and other processing in a cloud computing environment.
[0086] The above-described means and functions are realized by a computer (including a CPU, an information processing device, and various terminals) reading and executing a predetermined program. The program is provided, for example, in the form of a cloud service or SaaS (Software as a Service) provided from one or more computers via a network. The program is also provided, for example, in the form of a program recorded on a computer-readable recording medium. In this case, the computer reads the program from the recording medium, transfers it to an internal or external recording device, records it, and executes it. The program may also be pre-recorded on a recording device (recording medium) such as a magnetic disk, optical disk, or magneto-optical disk, and provided to the computer from the recording device via a communication line.
[0087] Although the embodiments of the present invention have been described above, the present invention is not limited to these embodiments. Furthermore, the effects described in the embodiments of the present invention are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments of the present invention. [Explanation of symbols]
[0088] 1 Image recognition system, 2 Computer, 21 Image generation unit, 22 Time stamping unit, 23 Question unit, 100 Frame image, 200 Still image data, 300 Prompt
Claims
1. An image recognition system that recognizes the situation of an object that changes over time, an image generating unit that generates still image data by combining two or more frame images extracted at different times from each video image obtained by photographing the object from one or more directions for a predetermined period of time; a time assigning unit that assigns a time to each of the two or more frame images included in the generated still image data in order to indicate the chronological order in which the two or more frame images were captured; a questioning unit that uses the still image data and a predetermined prompt to ask the generated AI a predetermined question about the situation and obtains an answer from the generated AI; An image recognition system comprising:
2. the predetermined prompt includes a role for answering, predetermined constraints, and an output format; The image recognition system according to claim 1 .
3. The still image data is extracted at predetermined intervals from a video of the object. The image recognition system according to claim 1 .
4. The predetermined constraints further include the obtained answer. The image recognition system according to claim 2 .
5. the image generation unit performs noise removal, resolution adjustment, or image normalization on the two or more frame images before generating the still image data; The image recognition system according to claim 1 .
6. The questioning unit asks the predetermined questions regarding the situation to a plurality of generation AIs, and obtains and integrates answers from the plurality of generation AIs. The image recognition system according to claim 1 .
7. When audio data, sensor data, or text data is associated with each of the two or more frame images, the image generation unit integrates the audio data, the sensor data, or the text data into the still image data. The image recognition system according to claim 1 .
8. The questioning unit outputs the still image data and the answer of the generated AI, and receives the predetermined question and feedback from the user. The image recognition system according to claim 1 .
9. the image generation unit selects the two or more frame images based on the received answer or feedback, and generates the still image data. The image recognition system according to claim 8 .
10. The system is executed in either an on-premise computing environment or a cloud computing environment, and the still image data and the answer of the generated AI are managed in the user's desired environment. The image recognition system according to any one of claims 1 to 9.
11. An image recognition method executed by a computer to recognize a time-series changing situation of an object, comprising: generating still image data by combining two or more frame images extracted at different times from each video image captured from one or more directions for a predetermined period of time; a step of assigning a time to each of the two or more frame images included in the generated still image data in order to indicate the chronological order in which the two or more frame images were captured; Using the still image data and a predetermined prompt, asking a generated AI a predetermined question about the situation and obtaining an answer from the generated AI; An image recognition method comprising:
12. On the computer, A step of generating still image data by combining two or more frame images extracted at different times from each video obtained by photographing an object from one or more directions for a predetermined period of time into a single image data; a step of assigning a time to each of the two or more frame images included in the generated still image data in order to indicate the chronological order in which the two or more frame images were captured; Using the still image data and a predetermined prompt, asking the generated AI a predetermined question about the situation and obtaining an answer from the generated AI; A computer-readable program for executing the program.
Citation Information
Patent Citations
Method and system for improving video question-answering precision based on multi-modal fusion model
CN112559698A
Method, computer program and device for performing artificial intelligence-based video question answering in data processing system (neural-symbolic action transformers for video question answering)
JP2023016740A
Machine learning device, machine learning method, machine learning program and inference device
JP2023117248A
Scene-aware video dialogue
JP2023510430A
Response generation device, response generation method, and response generation program
WO2023157265A1