Information processing system and information processing method

The system uses AI agents to generate and compare text information from reference and target images, addressing the challenge of complex condition settings in event detection, ensuring accurate event detection without them.

JP2026062056APending Publication Date: 2026-04-09SECOM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing information processing systems require complex condition settings to detect target events, making it difficult to accurately determine their occurrence.

Method used

An information processing system utilizing multiple AI agents to generate and compare text information from reference and target images, enabling accurate detection of events without complex condition settings.

Benefits of technology

Enables accurate detection of target events by generating and comparing text information from reference and target images, eliminating the need for complex condition settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062056000001_ABST
    Figure 2026062056000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing system and method that accurately detect whether or not a target event is occurring without requiring complex condition settings. [Solution] In the information processing system 1, the center device includes: a first generation unit that inputs a reference image of the monitoring area captured when no target event is occurring in the monitoring area to a first generation AI and causes the first generation AI to generate reference text information describing the context captured in the reference image; a second generation unit that inputs a target image of the monitoring area captured to a second generation AI and causes the second generation AI to generate target text information describing the context captured in the target image; a third generation unit that inputs information related to the reference text information, information related to the target text information, and information related to the target event to a third generation AI and causes the third generation AI to generate judgment result text information indicating whether or not the target event is included in the target image; and an output control unit that outputs information related to the judgment result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system and an information processing method.

Background Art

[0002] Conventionally, an information processing system has been developed in which an image of a monitoring area captured by a surveillance camera is visually monitored by a monitor.

[0003] In Patent Document 1, a notification device is disclosed that detects the posture of a first person from an image of a monitoring area, tracks the behavior of a second person, and determines that a notification is required when the posture of the first person matches a pre-stored posture and the behavior of the second person matches a pre-stored specific behavior.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the case of Patent Document 1, in order to detect a target event to be detected, complex condition settings such as postures and specific behaviors are required in advance. Therefore, it is required to accurately detect whether or not a target event has occurred without performing complex condition settings.

[0006] An object of the present invention is to provide an information processing system and an information processing method capable of accurately detecting whether or not a target event has occurred without performing complex condition settings.

Means for Solving the Problems

[0007] To solve these problems, the present invention provides an information processing system comprising: a first generation unit that inputs a reference image of a monitoring area captured when no target event is occurring in the monitoring area to a first generation AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generation AI or AI agent to generate reference text information describing the context captured in the reference image; a second generation unit that inputs a target image of the monitoring area captured to a second generation AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the second generation AI or AI agent to generate target text information describing the context captured in the target image; a third generation unit that inputs information related to reference text information, information related to target text information, and information related to target events to a third generation AI or AI agent that generates and outputs text information corresponding to a plurality of input text pieces, causing the third generation AI or AI agent to generate determination result text information indicating whether or not the target image contains a target event; and an output unit that outputs information related to the determination result text information.

[0008] In this information processing system, it is preferable that the third generation unit inputs information about reference text information and information about target text information to a third generation AI or AI agent, causes the third generation AI or AI agent to generate comparison result text information representing the comparison result between the information about reference text information and the information about target text information, inputs information about the comparison result text information and information about the target event to the third generation AI or AI agent, and causes the generation AI or AI agent to generate judgment result text information.

[0009] In this information processing system, it is preferable for the third generation unit to further cause a third generation AI or AI agent to generate, as determination result text information, information indicating the reason why it determined that the target image contains the target event.

[0010] In this information processing system, it is preferable that the third generation unit further inputs information indicating noteworthy points regarding the target event to the third generation AI or AI agent, causing the third generation AI or AI agent to generate judgment result text information.

[0011] In this information processing system, it is preferable that the first generation unit inputs multiple reference images to a first generation AI or AI agent and causes the first generation AI or AI agent to generate multiple reference text information, and the third generation unit uses the multiple reference text information to cause the third generation AI or AI agent to generate judgment result text information.

[0012] In this information processing system, it is preferable that the third generation unit generates a common portion of a plurality of reference text information using that plurality of reference text information, and then causes the third generation AI or AI agent to generate the judgment result text information based on that common portion.

[0013] In this information processing system, it is preferable that the first generation unit inputs multiple reference images to a first generation AI or AI agent, causes the first generation AI or AI agent to generate multiple reference text information, inputs the multiple reference text information to a fourth generation AI or AI agent that generates and outputs text information representing the common part of the multiple input text information, causes the fourth generation AI or AI agent to generate text information representing the common part of the multiple reference text information, and the third generation unit generates judgment result text information to a third generation AI or AI agent based on the common part.

[0014] In this information processing system, it is preferable that the first generation unit generates feature information generated by the first generation AI or AI agent when a reference image is input to the first generation AI or AI agent, and the third generation unit generates information related to reference text information based on the feature information.

[0015] In this information processing system, the third generation unit preferably identifies information about the reference text information based on the reference text information included in the group with the smallest difference from the target text information among the groups into which multiple reference text information have been classified.

[0016] In this information processing system, it is preferable to further include a reception unit that accepts the designation of target events by the monitor.

[0017] To solve these problems, the present invention provides an information processing method in which a reference image of a monitoring area captured when no target event has occurred in the monitoring area is input to a first generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generating AI or AI agent to generate reference text information describing the context captured in the reference image; a target image of the monitoring area captured is input to a second generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the second generating AI or AI agent to generate target text information describing the context captured in the target image; information regarding the reference text information, information regarding the target text information, and information regarding the target event are input to a third generating AI or AI agent that generates and outputs text information corresponding to a plurality of input text pieces of information, causing the third generating AI or AI agent to generate determination result text information indicating whether or not the target event is included in the target image, and outputting information regarding the determination result text information. [Effects of the Invention]

[0018] The information processing system and information processing method according to the present invention make it possible to accurately detect whether or not a target event is occurring without setting complex conditions. [Brief explanation of the drawing]

[0019] [Figure 1] This is a diagram showing the overall system configuration of Information Processing System 1. [Figure 2] It is a flowchart showing an example of the operation of the learning process. [Figure 3] It is a flowchart showing an example of the operation of the detection process. [Figure 4] (A) and (B) show an example of a reference image, and FIGS. 4(C) and (D) show an example of a target image. [Figure 5] It is a flowchart showing an example of the operation of other learning processes. [Figure 6] It is a flowchart showing an example of the operation of other detection processes.

Mode for Carrying Out the Invention

[0020] Hereinafter, the monitoring system according to the embodiment will be described with reference to the drawings.

[0021] FIG. 1 is a diagram showing the overall system configuration of the information processing system 1. As shown in FIG. 1, the information processing system 1 includes a monitoring device 10, a center device 20, a server device 30, etc. The number of each of the monitoring device 10, the center device 20, and the server device 30 is not limited to one, and may be plural. The monitoring device 10, the center device 20, and the server device 30 are connected to be mutually communicable via a network N. The network N is a wide area communication network such as an intranet or the Internet. The monitoring device 10 is installed in a monitoring target such as a house, a store, an office, a commercial facility, a factory, etc., and monitors one or a plurality of monitoring areas included in the monitoring target. The center device 20 is installed on a monitoring desk or the like in a monitoring center installed inside or outside the monitoring area, and aggregates and manages the monitoring results by the monitoring device 10. The server 30 is installed in a monitoring center or the like. The server 30 may be a cloud server in a wide area communication network such as the Internet.

[0022] The monitoring device 10 monitors for abnormalities such as intrusion, fire, and collapse of a person using sensors and other equipment installed on the target of monitoring. The monitoring device 10 includes an interface unit 11, an imaging unit 12, a monitoring sensor 13, a first input unit 14, a first output unit 15, a first communication unit 16, a first storage unit 17, and a first control unit 18.

[0023] The interface unit 11 has an interface circuit conforming to a serial bus standard such as USB, and communicates with the imaging unit 12 and the monitoring sensor 13 to send and receive various signals. Alternatively, the interface unit 11 may have an interface circuit conforming to a wired / wireless communication standard such as Ethernet®, IEEE 802.11, or Bluetooth® instead of an interface circuit conforming to a serial bus standard.

[0024] The imaging unit 12 is positioned in a monitoring area included in the object being monitored and images the monitoring area. The number of imaging units 12 in a single monitoring area is not limited to one, and there may be multiple units. The imaging unit 12 includes a camera. The camera has, for example, a photoelectric conversion element such as a CCD element or a C-MOS element, an imaging optical system that forms an image on the photoelectric conversion element, and an A / D converter that amplifies the electrical signal output from the photoelectric conversion element and performs analog-to-digital (A / D) conversion. The imaging unit 12 sequentially generates input images at a predetermined frame period and outputs them to the first control unit 18.

[0025] The monitoring sensor 13 is placed in the monitoring area included in the monitored object and monitors the monitoring area to detect abnormalities such as intrusion by suspicious persons, fire, or collapse of a person. The number of monitoring sensors 13 in one monitoring area is not limited to one, and there may be multiple sensors. For example, the monitoring sensor 13 is a magnetic sensor that detects intrusion by detecting the operation of opening and closing parts such as doors or windows of a building. The monitoring sensor 13 may also be an infrared sensor that detects intrusion by suspicious persons or collapse of a person by detecting the movement or state of moving objects inside the monitoring area based on changes in the amount of infrared light received. The monitoring sensor 13 may also be an ultrasonic sensor that emits ultrasonic waves and detects the movement or state of moving objects inside the monitoring area based on changes in the magnitude of the received ultrasonic waves to detect intrusion by suspicious persons or collapse of a person. The monitoring sensor 13 may also be a temperature sensor that detects intrusion by suspicious persons, fire, or collapse of a person by detecting the heat emitted by a person or object. The monitoring sensor 13 may be an image sensor that detects the movement or state of moving objects within the monitoring area based on the difference signal between the captured image and a pre-stored background image, thereby detecting the intrusion of a suspicious person or a person collapsing. The monitoring sensor 13 may also be an acceleration sensor or vibration sensor installed in a wearable device worn by a person in the monitoring area, and may detect a person falling based on the output information of the acceleration sensor or vibration sensor. When an abnormality is detected in the monitoring area, the monitoring sensor 13 transmits a detection signal to the first control unit 18.

[0026] The first input unit 14 has an interface circuit that receives signals from an operating device such as a keyboard, mouse, or touch panel, and accepts operations from the user, and outputs a signal corresponding to the accepted operation to the first control unit 18.

[0027] The first output unit 15 has a display device such as a liquid crystal display or an organic EL display and an interface circuit that outputs images to the display device, and displays various information such as images and text according to instructions from the first control unit 18. The first output unit 15 also has an audio output device such as a speaker and an interface circuit that outputs audio to the audio output device, and outputs audio according to instructions from the first control unit 18.

[0028] The first communication unit 16 has a communication interface circuit conforming to wired / wireless communication standards such as Ethernet® and IEEE 802.11, and communicates with the center device 20 via the network N to send and receive various information.

[0029] The first storage unit 17 includes semiconductor memory, magnetic storage media, and / or optical storage media. The first storage unit 17 stores the code of the computer program executed by the first control unit 18 to control the operation of the monitoring device 10, as well as various data. The computer program is installed in the first storage unit 17 by known methods via a computer-readable storage medium such as a CD-ROM or DVD-ROM, or via a communication line. The computer program may also be distributed from a server and installed in the first storage unit 17. The first storage unit 17 also stores as data the shooting locations captured by the imaging unit 12, i.e., the attributes of each monitoring area (house, store, office, commercial facility, factory, etc.).

[0030] The first control unit 18 includes a processor such as a CPU or multiprocessor and its peripheral circuits, and the processor controls the operation of the monitoring device 10 by executing a computer program stored in the first storage unit 17. A DSP, LSI, ASIC, FPGA, etc. may be used as the first control unit 18.

[0031] The central device 20 aggregates and displays the monitoring results from the monitoring device 10, thereby notifying the person being monitored. The central device 20 includes a second input unit 21, a second output unit 22, a second communication unit 23, a second storage unit 24, and a second control unit 25.

[0032] The second input unit 21 has an interface circuit that receives signals from an operating device such as a keyboard, mouse, or touch panel, and accepts operations from the user, and outputs a signal corresponding to the accepted operation to the second control unit 25.

[0033] The second output unit 22 has a display device such as a liquid crystal display or an organic EL display and an interface circuit that outputs images to the display device, and displays various information such as images and text according to instructions from the second control unit 25. The second output unit 22 also has an audio output device such as a speaker and an interface circuit that outputs audio to the audio output device, and outputs audio according to instructions from the second control unit 25.

[0034] The second communication unit 23 has a communication interface circuit that conforms to wired / wireless communication standards such as Ethernet (registered trademark) and IEEE 802.11, and communicates with the monitoring device 10 and the server device 30 via the network N to send and receive various information.

[0035] The second storage unit 24 includes semiconductor memory, magnetic storage media, and / or optical storage media. The second storage unit 24 stores the code and various data of the computer program executed by the second control unit 25 to control the operation of the center device 20. The computer program is installed in the second storage unit 24 by known methods via a computer-readable storage medium such as a CD-ROM or DVD-ROM, or via a communication line. The computer program may also be distributed from a server or the like and installed in the second storage unit 24.

[0036] The second control unit 25 includes a processor such as a CPU or multiprocessor and its peripheral circuits. The processor controls the operation of the center device 20 by executing a computer program stored in the second storage unit 24. A DSP, LSI, ASIC, FPGA, or the like may be used as the second control unit 25. The second control unit 25 includes, as a functional module of a program running on the processor, an acquisition unit 251, a first generation unit 252, a second generation unit 253, a third generation unit 254, an output control unit 255, and a receiving unit 256, etc.

[0037] The server device 30 stores one or more types of generating AI or AI agents, generates text information using the generating AI or AI agents in accordance with a request from the center device 20, and transmits it to the monitoring device 10. The server device 30 includes a third input unit 31, a third output unit 32, a third communication unit 33, a third storage unit 34, and a third control unit 35.

[0038] The third input unit 31 has an interface circuit that receives signals from an operating device such as a keyboard, mouse, or touch panel, and accepts operations from the user, and outputs a signal corresponding to the accepted operation to the third control unit 35.

[0039] The third output unit 32 has a display device such as a liquid crystal display or an organic EL display and an interface circuit for outputting images to the display device, and displays various information such as images and text according to instructions from the third control unit 35. The third output unit 32 also has an audio output device such as a speaker and an interface circuit for outputting audio to the audio output device, and outputs audio according to instructions from the third control unit 35.

[0040] The third communication unit 33 has a communication interface circuit conforming to wired / wireless communication standards such as Ethernet (registered trademark) and IEEE 802.11, and communicates with the center device 20 via the network N to send and receive various information.

[0041] The third storage unit 34 includes semiconductor memory, magnetic storage media, and / or optical storage media. The third storage unit 34 stores the code and various data of the computer program executed by the third control unit 35 to control the operation of the server device 30. The computer program is installed in the third storage unit 34 by known methods via a computer-readable storage medium such as a CD-ROM or DVD-ROM, or via a communication line. The computer program may also be distributed from a server or the like and installed in the third storage unit 34.

[0042] The third memory unit 34 stores one or more artificial intelligence (AIs). Each artificial intelligence is pre-trained to generate and output information corresponding to the input information. The generative AI includes one or more VLMs (Vision Language Models) that, when given an image and natural language (text) as input, generate and output information corresponding to the input image and natural language. This generative AI generates and outputs text information that describes the context (events, situations, content, objects, etc.) depicted in the input image. Each generative AI, which is a VLM, used in each process described later may be the same model or different models. Examples of VLMs include LLaVA (Large Language and Vision Assistant), GPT (Generative Pre-trained Transformer)-4v, and GPT-4o. LLaVA is a VLM that combines Llama2, an LLM (Large Language Model), with CLIP, which includes a vision encoder and a text encoder. LLaVA outputs an answer that is appropriate to an image, given a single image and natural language indicating a question about that image. LLaVA converts images into feature vectors using CLIP's vision encoder, then converts these feature vectors into a format that can be input to Llama2, and also converts natural language into feature vectors using a tokenizer. LLaVA inputs the feature vectors converted from the image and the feature vectors converted from the natural language into Llama2 and outputs a response that matches the image.

[0043] Furthermore, the generative AI includes one or more LLMs (Large-Scale Language Models) that, when natural language is input, generate and output information corresponding to the input natural language. This generative AI generates and outputs text information corresponding to multiple input text pieces. Alternatively, this generative AI generates and outputs text information representing the common part of multiple input text pieces. Each generative AI, which is an LLM used in each process described later, may be the same model or different models. Examples of LLMs include Llama2, ChatGPT, and BERT (Bidirectional Encoder Representations from Transformers).

[0044] The third control unit 35 includes a processor such as a CPU or multiprocessor and its peripheral circuits. The processor controls the operation of the server device 30 by executing a computer program stored in the third storage unit 34. A DSP, LSI, ASIC, FPGA, or the like may be used as the third control unit 35.

[0045] Figure 2 is a flowchart illustrating an example of the operation of the learning process by the center device 20. The learning process is a process for memorizing (learning) the state in which no target event has occurred in the monitoring area. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program that is pre-stored in the second storage unit 24.

[0046] First, the acquisition unit 251 acquires a reference image of the monitoring area when no target event is occurring in the monitoring area (step S101). The target event is an event that the monitor should detect and is an event that may occur in the monitoring area. Target events include abnormalities such as intrusion, fire, or a person collapsing. The target event can be any event, such as the presence or absence or movement of a predetermined object such as a car or luggage, or a fight. The reference image is a normal image of the monitoring area in a normal state. The reference image is, for example, an input image of the monitoring area captured by the imaging unit 12 when the monitoring sensor 13 has not detected any abnormality. The reference image may also be an input image captured by the imaging unit 12 when it has been confirmed that no target event has occurred, regardless of the detection status of the monitoring sensor 13. The reference image may be a still image or a video. The acquisition unit 251 transmits a reference image request signal to the monitoring device 10 via the second communication unit 23 to request the acquisition of a reference image. When the first control unit 18 of the monitoring device 10 receives a reference image request signal from the center device 20 via the first communication unit 16, it transmits a reference image to the center device 20 via the first communication unit 16. The acquisition unit 251 acquires the reference image by receiving it from the monitoring device 10 via the second communication unit 23. The first control unit 18 may also spontaneously transmit a reference image to the center device 20 each time the imaging unit 12 generates a reference image.

[0047] Next, the acquisition unit 251 uses known image processing techniques such as background subtraction or interframe subtraction to determine whether or not a change region exists in the acquired reference image (step S102). If no change region exists in the reference image, the acquisition unit 251 returns to step S101 and repeats the processing from steps S101 to S102.

[0048] On the other hand, if a region of change exists within the reference image, the first generation unit 252 obtains reference text information relating to the reference image from the reference image (step S103). For example, the first generation unit 252 obtains reference text information relating to the reference image from the reference image using a generation AI which is a VLM. This generation AI is an example of a first generation AI. The reference text information is, for example, information that shows text (natural language) that describes the context depicted in the reference image. The first generation unit 252 generates reference prompt information that includes instructions for generating reference text information from the reference image. The reference prompt information is text information that includes instructions for describing the context of the reference image. Preferably, the reference prompt information includes conditions for describing the reference image (conditions for specifying the text format to be output, conditions for points that require detailed explanation, conditions for points that should be ignored, etc.). The reference prompt information may be, for example, "Please describe the situation depicted in the image in detail using bullet points. Pay particular attention to the lighting conditions and the presence or absence of people."

[0049] The first generation unit 252 transmits a reference request signal to the server device 30 via the second communication unit 23, instructing the generation AI to generate reference text information by inputting a reference image and reference prompt information into the generation AI. The reference request signal includes the reference image and reference prompt information. When the third control unit 35 of the server device 30 receives a reference request signal from the center device 20 via the third communication unit 33, it inputs the reference image and reference prompt information included in the received reference request signal into the generation AI, which is a VLM, causing the generation AI to generate reference text information. The third control unit 35 transmits the generated reference text information to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the reference text information by receiving it from the server device 30 via the second communication unit 23. In this way, the first generation unit 252 inputs a reference image into the generation AI and causes the generation AI to generate reference text information related to the reference image. As will be described later, the learning process is executed repeatedly, and the first generation unit 252 inputs multiple reference images into the generation AI and causes the generation AI to generate multiple reference text information.

[0050] Next, the first generation unit 252 reads normal text information from the second storage unit 24 (step S104). Normal text information is information generated / stored based on one or more reference text information, and is stored in the second storage unit 24 in the processing described later. It is text information that indicates the context of the monitoring area when the target event has not occurred. Normal text information is stored in groups for multiple scenes according to the context of the monitoring area. For example, normal text information representing the scene when the lights in the monitoring area are off, normal text information representing the scene when the lights in the monitoring area are on and a person is present, and normal text information representing the scene when the lights in the monitoring area are on and no person is present are stored separately. Normal text information representing each scene is generated / stored based on one or more reference text information, as described later.

[0051] Next, the first generation unit 252 acquires difference text information that shows the difference between the acquired reference text information and the normal text information for each scene read from the second storage unit 24 (step S105). The first generation unit 252 acquires difference text information using, for example, a generation AI which is an LLM. For each normal text information of each scene, the first generation unit 252 generates difference prompt information that includes an instruction to generate the difference between the reference text information and the normal text information. The instruction includes the reference text information and the normal text information. The difference prompt information is, for example, "Please describe the difference between [reference text information] and [normal text information] in detail in bullet points." (The content shown in each text information is written in []).

[0052] The first generation unit 252 inputs the differential prompt information generated for each scene into the generation AI and sends a differential request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate differential text information for each scene. The differential request signal includes the differential prompt information for each scene. When the third control unit 35 of the server device 30 receives a differential request signal from the center device 20 via the third communication unit 33, it inputs the differential prompt information contained in the received differential request signal into the generation AI, which is an LLM, and causes the generation AI to generate differential text information. The third control unit 35 sends the generated differential text information for each scene to the center device 20 via the third communication unit 33. The first generation unit 252 acquires each differential text information by receiving it from the server device 30 via the second communication unit 23.

[0053] Next, the first generation unit 252 extracts the difference text information of the scene with the smallest difference between the reference text information and the normal text information from the acquired difference text information of each scene as the minimum difference text information (step S106). The first generation unit 252 extracts the minimum difference text information using a generation AI, such as an LLM. The first generation unit 252 generates minimum difference prompt information that includes an instruction to extract the difference text information of the scene with the smallest difference between the reference text information and the normal text information from the difference text information of each scene. The instruction includes the difference text information of each scene. The minimum difference prompt information is, for example, "Please select the one with the smallest difference from [Difference Text Information A], [Difference Text Information B]... However, please consider the presence or absence of people and lighting conditions as significant differences." (The content shown in each text information is written in []. Difference Text Information A indicates the difference text information of scene A, and Difference Text Information B indicates the difference text information of scene B.)

[0054] The first generation unit 252 sends a minimum difference request signal to the server device 30 via the second communication unit 23, instructing the generation AI to input minimum difference prompt information and extract minimum difference text information. The minimum difference request signal includes minimum difference prompt information. When the third control unit 35 of the server device 30 receives a minimum difference request signal from the center device 20 via the third communication unit 33, it inputs the minimum difference prompt information contained in the received minimum difference request signal to the generation AI, which is an LLM, causing the generation AI to extract minimum difference text information. The third control unit 35 transmits the extracted minimum difference text information to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the minimum difference text information by receiving it from the server device 30 via the second communication unit 23.

[0055] Next, the first generation unit 252 determines whether the extracted (acquired) minimum difference text information satisfies predetermined conditions (step S107). The predetermined conditions are, for example, that the items (differences) shown in the minimum difference text information are minor. The first generation unit 252 uses, for example, a generation AI which is LLM to determine whether the minimum difference text information satisfies predetermined conditions. The first generation unit 252 generates condition determination prompt information which includes an instruction to determine whether the minimum difference text information satisfies predetermined conditions. The instruction includes the minimum difference text information. The condition determination prompt information is, for example, "Is the difference shown in [minimum difference text information] minor? However, the presence or absence of people, the state of people, changes in lighting conditions, and dangerous conditions are not minor." (The content shown in each text information is described in []).

[0056] The first generation unit 252 inputs condition determination prompt information to the generation AI and sends a condition determination request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate a determination result. The condition determination request signal includes condition determination prompt information. When the third control unit 35 of the server device 30 receives a condition determination request signal from the center device 20 via the third communication unit 33, it inputs the condition determination prompt information included in the received condition determination request signal to the generation AI, which is an LLM, and causes the generation AI to generate a determination result. The third control unit 35 sends the generated determination result to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the determination result by receiving it from the server device 30 via the second communication unit 23. The first generation unit 252 determines whether the minimum difference text information satisfies the predetermined conditions based on whether the acquired determination result indicates that the minimum difference text information satisfies the predetermined conditions.

[0057] If the minimum difference text information does not meet the predetermined conditions, that is, if the difference between the reference text information and each of the already stored normal text information is not negligible, the first generation unit 252 stores the reference text information in the second storage unit 24 as normal text information representing a new scene (step S108). As a result, normal text information representing a new scene is added to the second storage unit 24 based on the newly acquired reference image. Next, the first generation unit 252 returns to step S101 and repeats the processing from step S101 onwards.

[0058] On the other hand, if the minimum difference text information satisfies a predetermined condition, that is, if the difference between the reference text information and the normal text information of any of the already stored scenes is minor, the first generation unit 252 considers the pre-detection text information to be information representing the context of that scene (hereinafter referred to as the "common scene"). The first generation unit 252 then extracts the common part between the reference text information and the normal text information of the common scene (step S109). The first generation unit 252 extracts the common part using a generation AI, for example, an LLM. This generation AI is an example of the fourth generation AI. The first generation unit 252 generates common part extraction prompt information that includes an instruction to extract the common part between the reference text information and the normal text information of the common scene. The instruction includes the reference text information and the normal text information corresponding to the minimum difference text information. The common part extraction prompt information is, for example, "Please extract the common part between [reference text information] and [normal text information of the common scene]." (The content shown in each text information is written in []).

[0059] The first generation unit 252 inputs common part extraction prompt information to the generation AI and sends a common part extraction request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate the common part between the reference text information and the normal text information of the common scene. The common part extraction request signal includes common part extraction prompt information. When the third control unit 35 of the server device 30 receives the common part extraction request signal from the center device 20 via the third communication unit 33, it inputs the common part extraction prompt information contained in the received common part extraction request signal to the generation AI, which is an LLM, and causes the generation AI to generate text information representing the common part between the reference text information and the normal text information of the common scene. The third control unit 35 transmits the extracted text information representing the common part to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the text information representing the common part between the reference text information and the normal text information of the common scene by receiving it from the server device 30 via the second communication unit 23.

[0060] The first generation unit 252 inputs information about multiple reference text information into the generation AI, causing the generation AI to generate text information representing the common part of the multiple reference text information. This allows the information processing system 1 to extract the common part of the multiple reference text information simply and with high accuracy. Alternatively, the first generation unit 252 may generate the common part by rule-based processing without having the generation AI generate it. For example, the first generation unit 252 may perform word or phrase-based matching comparisons on the multiple reference text information and extract the matched text portion as the common part.

[0061] Next, the first generation unit 252 updates the normal text information of the common scene using the acquired text information indicating the common part (step S110). As a result, the content based on the newly acquired reference image is reflected in the normal text information stored in the second storage unit 24.

[0062] In this way, the first generation unit 252 classifies the multiple reference text information into multiple groups, each represented by ordinary text information that represents one or more scenes. This allows the information processing system 1 to compare the target image of the monitoring area with each reference image for each classified group (scene) in the processing described later. Therefore, the information processing system 1 can determine with high accuracy whether or not the target image of the monitoring area has differences from one or more reference images corresponding to each scene.

[0063] Next, the first generation unit 252 returns to step S101 and repeats the processing from step S101 onward.

[0064] Figure 3 is a flowchart illustrating an example of the operation of the detection process by the center device 20. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program pre-stored in the second storage unit 24. The detection process is executed when the monitoring sensor 13 detects an abnormality or the like. However, it is not limited to this, and the detection process may be executed at any arbitrary timing, such as at predetermined time intervals. In addition, the monitoring device 10 may transmit the captured image to the center device 20 each time the imaging unit 12 captures an image, and the center device 20 may execute the detection process each time it receives the captured image.

[0065] First, the acquisition unit 251 acquires a target image of the monitoring area and the imaging status of the monitoring area at the time the target image was taken (step S201). The target image is an image used to determine whether or not a target event has occurred in the monitoring area, and is an input image of the monitoring area taken by the imaging unit 12 when it is unknown whether or not a target event has occurred in the monitoring area. The target image may be a still image or a video. The imaging status of the monitoring area is the shooting location and / or shooting time captured by the imaging unit 12. The acquisition unit 251 transmits a target image request signal to the monitoring device 10 via the second communication unit 23, requesting the acquisition of the target image and imaging status. When the first control unit 18 of the monitoring device 10 receives the target image request signal from the center device 20 via the first communication unit 16, it transmits the input image received from the imaging unit 12 as the target image, along with the shooting location or shooting time stored in the first storage unit 17, to the center device 20 via the first communication unit 16. The acquisition unit 251 acquires the target image and the imaging status of the monitoring area by receiving them from the monitoring device 10 via the second communication unit 23. The first control unit 18 may also spontaneously transmit the target image to the center device 20 whenever the imaging unit 12 generates a new input image when the monitoring sensor 13 has not detected any abnormality.

[0066] Next, the acquisition unit 251 uses known image processing techniques such as background subtraction or interframe subtraction to determine whether or not a change region exists in the acquired target image (step S202). If no change region exists in the target image, the acquisition unit 251 returns to step S201 and repeats the processing from steps S201 to S202.

[0067] On the other hand, if a region of change exists within the target image, the second generation unit 253 obtains target text information related to the target image from the target image (step S203). For example, the second generation unit 253 obtains target text information related to the target image from the target image using a generation AI which is a VLM. This generation AI is an example of a second generation AI. The target text information is, for example, information that shows text (natural language) that describes the context depicted in the target image. The second generation unit 253 generates target prompt information that includes instructions for generating target text information from the target image. The target prompt information is text information that includes instructions for describing the context of the target image. Preferably, the target prompt information includes conditions for describing the target image. For example, the target prompt information may be "Please describe the situation depicted in the image in detail using bullet points. Pay particular attention to the lighting conditions and the presence or absence of people."

[0068] The second generation unit 253 transmits a target request signal to the server device 30 via the second communication unit 23, instructing the generation AI to generate target text information by inputting the target image and target prompt information into the generation AI. The target request signal includes the target image and target prompt information. When the third control unit 35 of the server device 30 receives the target request signal from the center device 20 via the third communication unit 33, it inputs the target image and target prompt information contained in the received target request signal into the generation AI, which is a VLM, causing the generation AI to generate target text information. The third control unit 35 transmits the generated target text information to the center device 20 via the third communication unit 33. The second generation unit 253 acquires the target text information by receiving it from the server device 30 via the second communication unit 23. In this way, the second generation unit 253 inputs the target image into the generation AI and causes the generation AI to generate target text information related to the target image.

[0069] Next, the third generation unit 254 reads out each normal text information from the second storage unit 24 (step S204).

[0070] Next, the third generation unit 254 acquires differential text information that shows the difference between the acquired target text information and the normal text information for each scene read from the second storage unit 24 (step S205). The third generation unit 254 acquires differential text information using a generation AI, for example, an LLM. This generation AI is an example of a third generation AI. For each normal text information of each scene, the third generation unit 254 generates differential prompt information that includes an instruction to generate the difference between the target text information and the normal text information. The instruction includes the target text information and the normal text information. The differential prompt information is, for example, "Please describe the difference between [target text information] and [normal text information] in detail in bullet points." (The content shown in each text information is written in []).

[0071] The third generation unit 254 inputs the differential prompt information generated for each scene into the generation AI and sends a differential request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate differential text information for each scene. The differential request signal includes the differential prompt information for each scene. When the third control unit 35 of the server device 30 receives a differential request signal from the center device 20 via the third communication unit 33, it inputs the differential prompt information contained in the received differential request signal into the generation AI, which is an LLM, and causes the generation AI to generate differential text information. The third control unit 35 sends the generated differential text information for each scene to the center device 20 via the third communication unit 33. The third generation unit 254 acquires each differential text information by receiving it from the server device 30 via the second communication unit 23.

[0072] The content shown in the normal text information and the content shown in the reference text are examples of information related to the reference text information, and the content shown in the target text information is an example of information related to the target text information. The difference text information is an example of comparison result text information that represents the comparison result between the information related to the reference text information and the information related to the target text information. That is, the third generation unit 254 inputs information related to the reference text information, information related to the target text information, and information indicating instructions for generating comparison result text information to the generation AI, causing the generation AI to generate comparison result text information. As a result, the information processing system 1 can easily compare the state in the monitoring area where no target event has occurred with the state at the time of capturing the target image, and can accurately determine whether or not the target image contains the target event.

[0073] In particular, the third generation unit 254 generates information about reference text information using multiple reference text information, and causes the generation AI to generate comparison result text information that represents the comparison result between the information about the reference text information and the information about the target text information. By using multiple reference images and multiple reference texts, the information processing system 1 can compare using a variety of states in which the target event does not occur in the monitoring area, and can identify the target event included in the target image with high accuracy.

[0074] Furthermore, the third generation unit 254 generates information about reference text information based on the common parts of multiple reference text information, and causes the generation AI to generate comparison result text information that represents the comparison result between the information about the reference text information and the information about the target text information. By using the common parts of multiple reference text information, the information processing system 1 can perform a comparison based on common features of a state in which the target event does not occur in the monitoring area, and can identify the target event contained in the target image with higher accuracy.

[0075] Next, the third generation unit 254 extracts the difference text information of the scene with the smallest difference between the target text information and the normal text information from the acquired difference text information of each scene as the minimum difference text information (step S206). The third generation unit 254 extracts the minimum difference text information using a generation AI, such as an LLM. The third generation unit 254 generates minimum difference prompt information that includes an instruction to extract the difference text information with the smallest difference between the target text information and the normal text information from the difference text information of each scene. The instruction includes the difference text information of each scene. The minimum difference prompt information is, for example, "Please select the one with the smallest difference from [Difference Text Information A], [Difference Text Information B]... However, please consider the presence or absence of people and lighting conditions as significant differences." (The content shown in each text information is written in []. Difference Text Information A indicates the difference text information of scene A, and Difference Text Information B indicates the difference text information of scene B.)

[0076] The third generation unit 254 sends a minimum difference request signal to the server device 30 via the second communication unit 23, instructing the generation AI to input minimum difference prompt information and extract minimum difference text information. The minimum difference request signal includes minimum difference prompt information. When the third control unit 35 of the server device 30 receives a minimum difference request signal from the center device 20 via the third communication unit 33, it inputs the minimum difference prompt information contained in the received minimum difference request signal to the generation AI, which is an LLM, causing the generation AI to extract minimum difference text information. The third control unit 35 transmits the extracted minimum difference text information to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the minimum difference text information by receiving it from the server device 30 via the second communication unit 23.

[0077] The content shown in the normal text information corresponding to the minimum difference text information is an example of information related to the reference text information, and the content shown in the target text information is an example of information related to the target text information. Furthermore, the minimum difference text information is an example of comparison result text information. In other words, the third generation unit 254 inputs information related to the reference text information, information related to the target text information, and information indicating instructions for generating comparison result text information to the generation AI, causing the generation AI to generate the comparison result text information. As a result, the information processing system 1 can easily compare the state in the monitoring area when no target event has occurred with the state at the time the target image was captured.

[0078] Furthermore, the third generation unit 254 identifies the normal text information corresponding to the minimum difference text information based on the reference text information included in the group with the smallest difference from the target text information, from among the groups (scenes) represented by each normal text information, into which multiple reference text information have been classified. By comparing this normal text information corresponding to the minimum difference text information with the target text information, the third generation unit 254 can obtain a comparison result that focuses only on differences based on comparisons between the same scenes, without outputting differences based on comparisons between clearly different scenes (for example, differences in the scenes themselves such as "the brightness of the room is different") as comparison results, and can detect the target event in the target text information with high accuracy.

[0079] Next, the third generation unit 254 determines whether or not the target image contains the target event based on the extracted (acquired) minimum difference text information (step S207). The third generation unit 254 determines whether or not the target image contains the target event using a generation AI, for example, an LLM. The third generation unit 254 generates event determination prompt information that includes an instruction to determine whether or not the minimum difference text information indicates that the target image contains the target event. The instruction includes the minimum difference text information and information indicating the target event. The event determination prompt information may include instructions to specify the reason for determining whether or not the target event is included in the target image. That is, the event determination prompt information may include instructions to describe the reason for determining whether or not the minimum difference text information indicates that the target event is included in the target image. This allows the information processing system 1 to present information to the monitor to correctly understand the status of the monitoring area. Furthermore, the event determination prompt information may also include information regarding the imaging status of the monitoring area. This allows the information processing system 1 to determine with greater accuracy whether or not the target event is included in the target image. If the imaging status includes information that identifies the location of the monitoring area and information that identifies the time of shooting, the information processing system 1 can take these circumstances into consideration to determine with greater accuracy whether or not the target event has occurred. For example, for a person lying face down on a desk, the information processing system 1 can determine that the target event is "resting" if the shooting time is during a break, and that the target event is "feeling unwell" if it is outside of a break. Similarly, the information processing system 1 can determine that the target event is "resting" if the person is lying on a bed, and that the target event is "falling" if it is in a hallway. Furthermore, the event determination prompt information may include information indicating points of interest regarding the target event. Points of interest regarding the target event include, for example, objects that are highly likely to be related to the anomaly, such as people or cars, or locations where the anomaly is likely to occur, such as the vicinity of equipment. This allows the information processing system 1 to determine with greater accuracy whether or not the target image contains the target event. If the monitoring area is an office, the event determination prompt information may be something like, "Does the difference shown in [Minimum Difference Text Information] indicate an abnormal or dangerous sign in the office? At a minimum, fire, sudden illness, vandalism, etc., are considered dangerous signs, but not limited to those. Pay particular attention to the presence and condition of people when making your determination. Also, describe the reason why you determined that an abnormality occurred in the office or why you determined that no abnormality occurred." (The content shown in each text information is written in []). In this example, "abnormal or dangerous sign" and "At a minimum, fire, sudden illness, vandalism, etc., are considered dangerous signs, but not limited to those." are the information that indicates the target event.

[0080] The third generation unit 254 inputs event determination prompt information to the generation AI and sends an event determination request signal via the second communication unit 23 to the server device 30 to cause the generation AI to generate determination result text information indicating whether or not the target image contains the target event. The event determination request signal includes event determination prompt information. When the third control unit 35 of the server device 30 receives an event determination request signal from the center device 20 via the third communication unit 33, it inputs the event determination prompt information contained in the received event determination request signal to the generation AI, which is an LLM, and causes the generation AI to generate determination result text information. The third control unit 35 sends the generated determination result text information to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the determination result text information by receiving it from the server device 30 via the second communication unit 23. The third generation unit 254 determines whether or not the target image contains the target event based on whether or not the acquired determination result text information indicates that the target image contains the target event.

[0081] The minimum difference text information is an example of information regarding the reference text information and the target text information, and the information indicating the target event is an example of information regarding the target event. Specifically, the third generation unit 254 inputs information regarding the reference text information, information regarding the target text information, information regarding the target event, and information indicating instructions for generating information indicating the determination result of whether or not the target image contains the target event and / or the reason for determining that the target image contains the target event, causing the generation AI to generate the determination result text information, which is the determination result of whether or not the target image contains the target event and / or the reason for determining that the target image contains the target event. As a result, the information processing system 1 can appropriately determine whether or not the target image contains the target event or present the reason for determining that the target image contains the target event to the observer.

[0082] Furthermore, minimum difference text information is an example of information related to comparison result text information. Specifically, the third generation unit 254 inputs information related to comparison result text information, information related to the target event, and information indicating instructions for generating information showing the determination result of whether or not the target image contains the target event and / or the reason for determining that the target image contains the target event, causing the generation AI to generate information showing the determination result of whether or not the target image contains the target event and / or the reason for determining that the target image contains the target event. As a result, the information processing system 1 can appropriately determine whether or not the target image contains the target event or present the reason for determining that the target image contains the target event to the observer.

[0083] Furthermore, the third generation unit 254 inputs information regarding reference text information, information regarding target text information, information regarding target events, information regarding the imaging status of the monitoring area, and instructions to generate information indicating the determination result of whether or not the target event is included in the target image and / or the reason for determining that the target event is included in the target image, causing the generation AI to generate information indicating the determination result of whether or not the target event is included in the target image and / or the reason for determining that the target event is included in the target image, taking into account the imaging status of the monitoring area. As a result, the information processing system 1 can appropriately determine whether or not the target event is included in the target image or present the reason for determining that the target event is included in the target image to the monitor.

[0084] Furthermore, the third generation unit 254 inputs information indicating noteworthy aspects of the target event to the generation AI, causing the generation AI to generate a determination result regarding whether or not the target event is included in the target image. As a result, the information processing system 1 can determine with greater accuracy whether or not the target event is included in the target image.

[0085] Furthermore, the third generation unit 254 causes the generating AI to generate a determination result of whether or not the target image contains the target event, based on the comparison result text information generated based on multiple reference text information. As a result, the information processing system 1 can determine with high accuracy whether or not the target image contains the target event.

[0086] Furthermore, the third generation unit 254 causes the generating AI to generate a determination result of whether or not the target image contains the target event, based on the comparison result text information generated based on the common parts of multiple reference text information. As a result, the information processing system 1 can determine with higher accuracy whether or not the target image contains the target event.

[0087] If the target image contains the target event, the output control unit 255 outputs the acquired determination result, i.e., that the target image contains the target event, and / or the reason for determining that the target image contains the target event, by displaying or outputting it as sound on the second output unit 22 (step S208). The output control unit 255 may also output predetermined information by adding or subtracting it from the statement that the target image contains the target event and / or the reason for determining that the target image contains the target event. This allows the monitor to easily and appropriately identify whether or not the target image contains the target event and / or the reason for determining that the target image contains the target event, and the information processing system 1 can monitor the monitoring area more reliably.

[0088] The fact that the target image contains the target event, and / or the reason for determining that the target image contains the target event, is an example of information regarding the determination result of whether or not the target image contains the target event. The output control unit 255 may also output the comparison result text information itself. This allows the monitor to appropriately identify the event contained in the target image, and the information processing system 1 to monitor the monitoring area more reliably.

[0089] Next, the third generation unit 254 returns to step S201 and repeats the processing from step S201 onward.

[0090] On the other hand, if the target image does not contain the target event, the third generation unit 254 extracts the common portion between the target text information and the normal text information corresponding to the extracted minimum difference text information (step S209). The third generation unit 254 extracts the common portion using a generation AI, for example, an LLM. The third generation unit 254 generates common portion extraction prompt information that includes an instruction to extract the common portion between the target text information and the normal text information corresponding to the minimum difference text information. The instruction includes the target text information and the normal text information corresponding to the minimum difference text information. The common portion extraction prompt information is, for example, "Please extract the common portion of [target text information] and [normal text information]." (The content shown in each text information is written in []).

[0091] The third generation unit 254 inputs common part extraction prompt information to the generation AI and sends a common part extraction request signal to the server device 30 via the second communication unit 23, causing the generation AI to extract the common part between the reference text information and the normal text information corresponding to the minimum difference text information. The common part extraction request signal includes common part extraction prompt information. When the third control unit 35 of the server device 30 receives the common part extraction request signal from the center device 20 via the third communication unit 33, it inputs the common part extraction prompt information included in the received common part extraction request signal to the generation AI, which is an LLM, causing the generation AI to extract the common part between the target text information and the normal text information corresponding to the minimum difference text information. The third control unit 35 transmits the extracted common part to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the common part between the target text information and the normal text information corresponding to the minimum difference text information by receiving it from the server device 30 via the second communication unit 23.

[0092] Next, the third generation unit 254 updates the normal text information of the scene corresponding to the minimum difference text information with text information indicating the acquired common part (step S210). As a result, the normal text information stored in the second storage unit 24 reflects the content based on the target image that does not contain the target event. Next, the third generation unit 254 returns to step S201 and repeats the processing from step S201 onwards.

[0093] Note that the process in step S102 in Figure 2 and / or the process in step S202 in Figure 3 may be omitted. Also, in the process in step S102 in Figure 2 and / or the process in step S202 in Figure 3, the acquisition unit 251 may use known image processing techniques to detect a predetermined object such as a person or a car from the reference image or target image, and proceed to step S103 or S203 only if a predetermined object is detected.

[0094] Furthermore, if it is determined in step S207 of Figure 3 that the target image does not contain the target event, the output control unit 255 may also output the determination result, that is, that the target image does not contain the target event, and / or the reason for determining that the target image does not contain the target event.

[0095] Furthermore, the process in step S109 in Figure 2 and / or the process in step S209 in Figure 3 may be omitted, and the normal text information may be updated with the latest reference text information or target text information. Also, the processes in steps S109 to S110 in Figure 2 and / or the processes in steps S209 to S210 in Figure 3 may be omitted, and the normal text information may not be updated. Also, the processes in steps S104 to S110 in Figure 2 and / or the processes in steps S204 to S206 and S209 to S210 in Figure 3 may be omitted. In that case, the third generation unit 254 acquires differential text information showing the difference between the target text information and the reference text information acquired immediately before, in the same manner as the process in step S205, and determines whether or not a target event has occurred in the monitoring area based on the acquired differential text information, in the same manner as the process in step S207.

[0096] Furthermore, the prompt information used in the learning and detection processes may not be pre-set but may be specified by the monitor. In this case, the reception unit 256 accepts the specification of each prompt information entered by the monitor using the second input unit 21 at any time. The reception unit 256 may also accept the specification of target events by the monitor. In this case, the reception unit 256 accepts the specification of target events entered by the monitor using the second input unit 21 at any time. Through these measures, the monitor can flexibly set the events they wish to detect in the monitoring area, and the information processing system 1 can improve the convenience of the monitor and enhance the security of the monitoring area.

[0097] Figures 4(A) and (B) show examples of reference images, and Figures 4(C) and (D) show examples of target images.

[0098] The reference images shown in Figures 4(A) and (B), and the target images shown in Figures 4(C) and (D), are images of the office taken at different times. In the reference images in Figures 4(A) and (B), and the target image in Figure 4(C), person P is working, but in the target image in Figure 4(D), person P is lying down.

[0099] The reference text information generated from the reference image in Figure 4(A) describes "Office scene: There is one person. He is wearing a white shirt and blue pants and is working at a computer. There is a computer on the desk. Lighting: It appears to be typical office lighting, and it is presumed that no special lighting is being used. Fluorescent lights and incandescent bulbs installed on the ceiling are likely being used. Floor: It appears to be a normal office floor, plain and simple in design." On the other hand, the reference text information generated from the reference image in Figure 4(B) describes "An office scene is depicted. In the center is a desk with three black monitors and a laptop computer. Person: One person is sitting and working with deep concentration (using a laptop computer). Lighting: The lighting setup is unknown. Wall: Wallpaper with a white base." The normal text information, which is a common part extracted from each reference text information, describes "There is a computer on the desk. One person is working (using a laptop computer)." Thus, although the reference images in Figures 4(A) and 4(B) are similar to each other, the content described in the reference text information generated from each reference image differs slightly. By extracting common parts from each reference image, the information processing system 1 can appropriately extract the characteristics of the monitoring area when no abnormalities have occurred.

[0100] The target text information generated from the target image in Figure 4(C) states: "Situation: A man is sitting at a desk and working on a computer. Presence of people: One man is visible. Three computer screens are present, and a laptop is open in front of the man. Lighting: The entire room relies on natural light. Since no windows are visible, it is unknown whether there is external light." The difference text information, which is the difference between the above normal text information and this target text information, states: "The laptop is open in front of the man. Lighting: The entire room relies on natural light." The result of the determination of whether or not an abnormality has occurred in the monitoring area states: "No abnormality. Reason: The laptop being open in front of the man is a common sight during work and is not unusual. The entire room relies on natural light, which is a normal office lighting environment and is not unusual." Based on these determination results, the monitor can easily understand that the monitoring sensor 13 had made a false detection.

[0101] The target text information generated from the target image in Figure 4(D) states: "Location: Office. Presence of people: 1 person (male). Condition: The man is lying on the floor. Clothing: Wearing a white shirt and black pants. Surrounding environment: The man is located near a desk, with a computer screen and laptop on the desk. Lighting conditions: The office has standard lighting settings, with no particularly noticeable lighting fixtures or special lighting effects." The difference text information, which is the difference between the above normal text information and this target text information, states: "A person is lying on the floor." The result of the determination of whether or not an abnormality has occurred in the monitoring area states: "Abnormality detected. Reason: It is unnatural for a person to be lying on the floor in a normal office." Based on these determination results, the monitor can easily understand that the cause detected by the monitoring sensor 13 was a person collapsing.

[0102] In this way, when the information processing system 1 detects a difference between the reference image and the target image, it uses the generating AI to notify the monitor whether or not an anomaly has occurred in the monitoring area and / or the reason for that determination. As a result, the monitor can easily and accurately grasp the situation in the monitoring area, regardless of their skills, experience, or prior knowledge of the monitoring area, and can respond appropriately to any anomalies occurring in the monitoring area.

[0103] As explained above, the information processing system 1 outputs text information indicating whether or not a target event is included in a target image, based on reference text information generated from a reference image by the generating AI and target text information generated from a target image by the generating AI. The information processing system 1 can appropriately determine whether or not a target event is included in a target image by utilizing the knowledge that the generating AI has learned. Therefore, the information processing system 1 can accurately detect whether or not a target event is occurring without setting complex conditions.

[0104] Figures 5 and 6 are flowcharts illustrating other embodiments.

[0105] Figure 5 is a flowchart showing an example of the operation of the learning process according to another embodiment. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program that is pre-stored in the second storage unit 24. Since the processing in steps S301 to S302 in Figure 5 is the same as the processing in steps S101 to S102 in Figure 2, only the processing in steps S303 to S304 will be described below.

[0106] If a change region exists within the reference image in step S302, the first generation unit 252 acquires reference feature information related to the reference image from the reference image (step S303). In the same manner as in step S103, the first generation unit 252 transmits a reference request signal to the server device 30 via the second communication unit 23, and the third control unit 35 of the server device 30 causes the generation AI to generate reference text information. However, the third control unit 35 acquires the feature quantities of the reference image, which are intermediate products generated by the generation AI, as reference feature information and transmits them to the center device 20 via the third communication unit 33. The intermediate products are obtained from the feature quantities output from the vision encoder of the VLM, which is a generation AI that generates pre-detection text information. The third control unit 35 acquires the feature quantities that the text decoder of the VLM can recognize as reference feature information. The first generation unit 252 acquires the reference feature information by receiving it from the server device 30 via the second communication unit 23. In this way, when the first generation unit 252 receives a reference image as input to the generation AI, it generates feature information generated by the generation AI.

[0107] Next, the first generation unit 252 stores the acquired reference feature information as normal feature information in the second storage unit 24 (step S304). Then, the first generation unit 252 returns to step S301 and repeats the process from step S301 onward. As this process is repeated, normal feature information consisting of multiple reference feature information is stored in the second storage unit 24.

[0108] Figure 6 is a flowchart showing an example of the operation of the detection process according to another embodiment. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program that is pre-stored in the second storage unit 24. Since the processing of steps S401 to S402 in Figure 6 is the same as the processing of steps S201 to S202 in Figure 3, only the processing of steps S403 to S410 will be described below.

[0109] If a change region exists within the target image in step S402, the second generation unit 253 acquires target feature information related to the target image from the target image (step S403). In the same manner as in step S203, the second generation unit 253 transmits a target request signal to the server device 30 via the second communication unit 23, and the third control unit 35 of the server device 30 causes the generation AI to generate target text information. However, the third control unit 35 acquires the feature quantities of the target image, which are intermediate products generated by the generation AI, as target feature information and transmits them to the center device 20 via the third communication unit 33. The third control unit 35 acquires the feature quantities that can be recognized by the text decoder of the generation AI, VLM, as target feature information. The second generation unit 253 acquires the target feature information by receiving it from the server device 30 via the second communication unit 23.

[0110] Next, the third generation unit 254 classifies the normal feature information, which consists of multiple pre-detection feature information stored in the second storage unit 24, into multiple groups by clustering (step S404). The third generation unit 254 clusters each normal feature information using a known statistical method such as the k-means method. Since the features are classified more flexibly and appropriately than text, the information processing system 1 can appropriately classify the features of images of monitoring areas where the target event has not occurred, according to environmental conditions.

[0111] Next, the third generation unit 254 identifies the group of normal feature information that best approximates the target feature information obtained in step S403 (step S405). For example, the third generation unit 254 calculates the centroid position of each group in the feature space and identifies the group with the centroid position having the smallest Euclidean distance from the target feature information as the group that best approximates the target feature information.

[0112] Next, the third generation unit 254 acquires the closest-to-the-nearest text information corresponding to the identified group and the target text information corresponding to the target feature information (step S406). For example, the third generation unit 254 calculates the feature corresponding to the centroid position of the identified group as the closest-to-the-nearest feature information. The third generation unit 254 sends a decode request signal via the second communication unit 23 to the server device 30 to cause the generation AI to generate (decode) the text information corresponding to the closest-to-the-nearest feature information and the text information corresponding to the target feature information. The decode request signal includes the closest-to-the-nearest feature information and the target feature information. When the third control unit 35 of the server device 30 receives the decode request signal from the center device 20 via the third communication unit 33, it inputs the closest-to-the-nearest feature information contained in the received decode request signal to the generation AI (VLM text decoder), which is an LLM, and causes the generation AI to generate the closest-to-the-nearest text information corresponding to the closest-to-the-nearest feature information. Furthermore, the third control unit 35 inputs the target feature information contained in the received decode request signal to the generating AI (VLM text decoder), which is an LLM, and causes the generating AI to generate target text information corresponding to the target feature information. The third control unit 35 transmits the generated similar text information and target text information to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the similar text information and target text information by receiving them from the server device 30 via the second communication unit 23.

[0113] Next, the third generation unit 254 determines whether or not the target image contains the target event based on the acquired similar text information and target text information (step S407). The third generation unit 254 determines whether or not the target image contains the target event using a generation AI, for example, LLM. The third generation unit 254 generates event determination prompt information and obtains the determination result in the same manner as the process in step S207 of Figure 3. However, the event determination prompt information in this embodiment includes an instruction to determine whether or not the target text information indicates that the target image contains the target event. The instruction includes similar text information and target text information. The instruction may also include information regarding the imaging conditions. If the monitoring area is an office, the event determination prompt information may be something like: "Does the difference shown in [Recently Similar Text Information] and [Target Text Information] indicate an abnormal or dangerous sign in the office? At a minimum, fire, sudden illness, vandalism, etc., should be considered dangerous signs, but not limited to these. Pay particular attention to the presence and condition of people when making your determination. Please also describe the reason why you determined that an abnormality occurred in the office or why you determined that no abnormality occurred." (The content shown in each text information is written in []).

[0114] Nearest similar text information is an example of information related to reference text information. That is, when the third generation unit 254 inputs a reference image to the generation AI, it generates information related to reference text information based on the feature information generated by the generation AI. In this case as well, the information processing system 1 can appropriately determine whether or not the target image contains the target event.

[0115] If an abnormality occurs in the monitoring area, the output control unit 255 outputs the acquired judgment result in the same manner as the process in step S208 (step S408). Next, the third generation unit 254 returns to step S401 and repeats the process from step S401 onwards.

[0116] On the other hand, if no abnormality occurs in the monitoring area, the third generation unit 254 stores the target feature information as normal feature information in the second storage unit 24 (step S409). Next, the third generation unit 254 returns to step S401 and repeats the processing from step S401 onwards.

[0117] As explained above, even when using feature information generated by the generating AI, the information processing system 1 can accurately detect whether or not a target event is occurring without setting complex conditions. In the other embodiments described above, the third generation unit 254 grouped the normal feature information in the detection process, but instead, the normal feature information may be grouped in the learning process. That is, the third generation unit 254 may group the normal features for each scene by performing a clustering process similar to the process in step S404 at the end of the learning process, and then extract the closest similarity group from each group (scene) in step S405 of the detection process.

[0118] Although preferred embodiments have been described above, the embodiments are not limited to the examples described above. For example, the information processing system 1 may have an AI agent generate each piece of information instead of a generating AI. In that case, the third storage unit 34 stores one or more AI agents. Each AI agent is pre-trained to generate and output information corresponding to the input information.

[0119] The AI ​​agent includes one or more autonomous AI agents using VLM that generate and output information corresponding to the input image and natural language when an image and natural language are input. This AI agent generates and outputs text information that describes the context depicted in the input image. The AI ​​agent also includes one or more autonomous AI agents using LLM that generate and output information corresponding to the input natural language when natural language is input. This AI agent generates and outputs text information corresponding to multiple input text pieces. Alternatively, this AI agent generates and outputs text information representing the common part of multiple input text pieces. When given a goal, the autonomous AI agent outputs information that achieves the goal by having a generating AI generate tasks to achieve that goal, collecting information to allow the generating AI to execute the generated tasks, and repeating the process of having the generating AI execute the tasks. The AI ​​agent may also autonomously acquire environmental information such as temperature and lighting of the monitored object (monitoring area), and output information that achieves the goal while considering the acquired environmental information.

[0120] In each process where a VLM-based generative AI is used, as shown in Figures 2, 3, 5, and 6, an AI agent using VLM is used, and in each process where an LLM-based generative AI is used, an AI agent using LLM is used. Each AI agent used in place of the 1st to 4th generative AIs is an example of the 1st to 4th AI agents, respectively. In each process, no prompt information is generated; instead, an AI agent generated to achieve the goal corresponding to each prompt information is used. Each VLM-based AI agent used in each process may be the same agent or different agents. Each LLM-based AI agent used in each process may be the same agent or different agents.

[0121] Information processing system 1 can accurately detect whether or not a target event is occurring, even when using an AI agent, without requiring complex condition settings.

[0122] Furthermore, in the information processing system 1, the learning process and detection process may be performed by the monitoring device 10 instead of the center device 20. In that case, the first control unit 18 of the monitoring device 10 has the parts of the second control unit 25 and performs the learning process and detection process. Alternatively, the learning process and detection process may be performed by the server device 30 instead of the center device 20. In that case, the third control unit 35 of the server device 30 has the parts of the second control unit 25 and performs the learning process and detection process. In that case, in step S208 in Figure 3 and step S408 in Figure 6, the server device 30 may transmit the determination result to the center device 20 or the monitoring device 10 via the third communication unit 33 and output it to the second output unit 22 or the first output unit 15.

[0123] Furthermore, each generating AI and / or each AI agent may be stored in the storage unit of the device that performs the learning and detection processing, rather than in the third storage unit 34 of the server device 30. In that case, when the device that performs the learning and detection processing uses each generating AI and / or each AI agent, it does not send each request signal to the server device 30, but instead causes each generating AI and / or each AI agent stored in the storage unit of the device to generate the information.

[0124] Furthermore, in the above embodiment, the first generation unit 252 generates a common portion based on a plurality of reference text information it has generated, and the third generation unit 254 generates the judgment result text information based on that common portion. However, the third generation unit 254 may generate the common portion using a plurality of reference text information generated by the first generation unit 252, and then generate the judgment result text information. In this case, the third generation unit 254 may generate the common portion using a fourth generation AI or AI agent, similar to the first generation unit 252. Alternatively, the third generation unit 254 may generate the common portion without using a fourth generation AI or AI agent. [Explanation of Symbols]

[0125] 1 Information processing system, 10 Monitoring device, 20 Center device, 30 Server device, 251 Acquisition unit, 252 First generation unit, 253 Second generation unit, 254 Third generation unit, 255 Output control unit, 256 Reception unit

Claims

1. A first generation unit inputs a reference image of the monitoring area, captured when no target event is occurring in the monitoring area, to a first generation AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generation AI or AI agent to generate reference text information describing the context captured in the reference image. The system inputs the target image captured in the aforementioned monitoring area to a second generating AI or AI agent that generates and outputs text information describing the context captured in the input image, and causes the second generating AI or AI agent to generate the target text information describing the context captured in the target image. A third generation unit inputs information relating to the reference text information, information relating to the target text information, and information relating to the target event to a third generation AI or AI agent that generates and outputs text information corresponding to a plurality of input text information, and causes the third generation AI or AI agent to generate determination result text information indicating whether or not the target event is included in the target image. An output unit that outputs information related to the judgment result text information, An information processing system characterized by having the following features.

2. The third generation unit is, The information relating to the reference text information and the information relating to the target text information are input to the third generating AI or AI agent, and the third generating AI or AI agent generates comparison result text information representing the comparison result of the information relating to the reference text information and the information relating to the target text information. The information processing system according to claim 1, wherein information relating to the comparison result text information and information relating to the target event are input to the third generating AI or AI agent, and the generating AI or AI agent generates the judgment result text information.

3. The information processing system according to claim 2, wherein the third generation unit causes the third generation AI or AI agent to further generate information indicating the reason for determining that the target image contains the target event, as determination result text information.

4. The information processing system according to any one of claims 1 to 3, wherein the third generation unit further inputs information indicating matters of note regarding the target event to the third generation AI or AI agent, and causes the third generation AI or AI agent to generate the determination result text information.

5. The first generation unit inputs a plurality of reference images to the first generation AI or AI agent, and causes the first generation AI or AI agent to generate a plurality of reference text information. The information processing system according to claim 1, wherein the third generation unit causes the third generation AI or AI agent to generate the determination result text information using the plurality of reference text information.

6. The information processing system according to claim 5, wherein the third generation unit generates a common portion of the plurality of reference text information using the plurality of reference text information, and causes the third generation AI or AI agent to generate the determination result text information based on the common portion.

7. The first generation unit inputs a plurality of reference images to the first generation AI or AI agent, causes the first generation AI or AI agent to generate a plurality of reference text information, inputs the plurality of reference text information to a fourth generation AI or AI agent that generates and outputs text information representing the common part of the plurality of input text information, causes the fourth generation AI or AI agent to generate text information representing the common part of the plurality of reference text information, The information processing system according to claim 1, wherein the third generation unit causes the third generation AI or AI agent to generate the determination result text information based on the common part.

8. When the first generation unit receives the reference image as input to the first generation AI or AI agent, it generates feature information generated by the first generation AI or AI agent. The information processing system according to any one of claims 1 to 3, wherein the third generation unit generates information relating to the reference text information based on the feature information.

9. The information processing system according to claim 5, wherein the third generation unit identifies information relating to the reference text information based on the reference text information included in the group with the smallest difference from the target text information among the groups into which the plurality of reference text information have been classified.

10. The information processing system according to claim 1, further comprising a reception unit for receiving designations of target events by a monitor.

11. A reference image of the monitoring area, captured when no target event occurs in the monitoring area, is input to a first generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generating AI or AI agent to generate reference text information describing the context captured in the reference image. The target image captured in the aforementioned monitoring area is input to a second generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the second generating AI or AI agent to generate the target text information describing the context captured in the target image. The information relating to the reference text information, the information relating to the target text information, and the information relating to the target event are input to a third generating AI or AI agent that generates and outputs text information corresponding to the multiple input text information, and the third generating AI or AI agent generates determination result text information indicating whether or not the target image contains the target event. Outputs information related to the judgment result text information. An information processing method characterized by the following:

Citation Information

Patent Citations

  • Notification device

    JP2012003597A