Information processing system and information processing method

The system uses AI agents to generate and compare pre- and post-detection text information, enhancing the accuracy of event detection and analysis in monitored areas by analyzing sensor images.

JP2026062046APending Publication Date: 2026-04-09SECOM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing information processing systems struggle to accurately monitor and analyze changes in monitored areas, particularly in detecting events and understanding the context of images captured by sensors.

Method used

An information processing system utilizing AI agents to generate and compare pre-detection and post-detection text information, enabling accurate analysis of monitoring area situations by inputting pre-detection images to a first generation AI for text generation, post-detection images to a second generation AI, and using a third generation AI to compare and output results.

Benefits of technology

Enhances the ability to accurately grasp the situation in monitored areas by identifying events and their causes through AI-generated text comparisons, improving event detection and analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062046000001_ABST
    Figure 2026062046000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing system and method that enable monitors to more accurately grasp the situation in their monitoring area. [Solution] The information processing system includes: a first generation unit that inputs a pre-detection image of the monitoring area, captured before detection by a sensor monitoring the monitoring area, to a first generation AI and causes the first generation AI to generate pre-detection text information describing the context captured in the pre-detection image; a second generation unit that inputs a post-detection image of the monitoring area, captured after detection by the sensor, to a second generation AI and causes the second generation AI to generate post-detection text information describing the context captured in the post-detection image; and a third generation unit that inputs information related to the pre-detection text information and information related to the post-detection text information to a third generation AI and causes the third generation AI to generate text information representing a comparison result of the information related to the pre-detection text information and the information related to the post-detection text information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system and an information processing method.

Background Art

[0002] Conventionally, an information processing system has been developed that monitors a monitoring area using sensors such as an infrared sensor, an ultrasonic sensor, and a temperature sensor, and allows a monitor to visually monitor an image of the monitoring area captured by a monitoring camera.

[0003] Patent Document 1 discloses a monitoring device in which an imaging unit that images a monitoring area and a detection unit that detects an intruder and transmits an abnormal signal are connected, and further connected to a monitoring center device via a communication line. When this monitoring device receives an abnormal signal from the detection unit, it transmits the image data captured by the imaging unit to the monitoring center device.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In an information processing system in which a monitor visually monitors an image of a monitored area, it is required that the monitor can more accurately grasp the situation of the monitored area.

[0006] An object of the present invention is to provide an information processing system and an information processing method that enable a monitor to more accurately grasp the situation of a monitored area.

Means for Solving the Problems

[0007] <00oooo36>To solve these problems, the present invention provides an information processing system comprising: a first generation unit that inputs a pre-detection image, captured of the monitoring area before detection by a sensor monitoring the monitoring area, to a first generation AI or AI agent that generates and outputs text information describing the context in the input image, causing the first generation AI or AI agent to generate pre-detection text information describing the context in the pre-detection image; a second generation unit that inputs a post-detection image, captured of the monitoring area after detection by the sensor, to a second generation AI or AI agent that generates and outputs text information describing the context in the input image, causing the second generation AI or AI agent to generate post-detection text information describing the context in the post-detection image; a third generation unit that inputs information regarding pre-detection text information and information regarding post-detection text information to a third generation AI or AI agent that generates and outputs text information as a result of comparing a plurality of input text information, causing the third generation AI or AI agent to generate comparison result text information regarding the pre-detection text information and the post-detection text information; and an output unit that outputs information regarding the comparison result text information.

[0008] In this information processing system, the third generation unit preferably inputs information about pre-detection text information and information about post-detection text information to a third generation AI or AI agent, causing the third generation AI or AI agent to generate information indicating whether or not an event targeted by the sensor has occurred in the monitoring area or the reason why the sensor detected the event, and the output unit preferably outputs information regarding whether or not an event has occurred in the monitoring area or the reason for detection.

[0009] In this information processing system, the third generation unit preferably inputs information regarding pre-detection text information, information regarding post-detection text information, and information regarding the imaging status of the monitoring area to a third generation AI or AI agent, and causes the third generation AI or AI agent to generate information indicating whether or not an event targeted by the sensor has occurred in the monitoring area, or the reason why the sensor detected the event, taking into account the imaging status, and the output unit preferably outputs information regarding whether or not an event has occurred in the monitoring area or the reason for detection.

[0010] In this information processing system, the pre-detection image is an image taken when no event to be detected by the sensor has occurred in the monitoring area. The first generation unit inputs multiple pre-detection images to a first generation AI or AI agent and causes the first generation AI or AI agent to generate multiple pre-detection text information. The third generation unit uses the multiple pre-detection text information to cause the third generation AI or AI agent to generate comparison result text information.

[0011] In this information processing system, it is preferable for the third generation unit to use multiple pre-detection text information to generate a common portion of the multiple pre-detection text information, and then have the third generation AI or AI agent generate comparison result text information based on that common portion.

[0012] In this information processing system, it is preferable that the first generation unit inputs multiple pre-detection images to a first generation AI or AI agent, causes the first generation AI or AI agent to generate multiple pre-detection text information, inputs the multiple pre-detection text information to a fourth generation AI or AI agent that generates and outputs text information representing the common part of the multiple input text information, causes the fourth generation AI or AI agent to generate text information representing the common part of the multiple pre-detection text information, and the third generation unit generates comparison result text information to a third generation AI or AI agent based on the common part.

[0013] In this information processing system, it is preferable that the first generation unit generates feature information generated by the first generation AI or AI agent when the pre-detection image is input to the first generation AI or AI agent, and the third generation unit generates information related to the pre-detection text information based on the feature information.

[0014] To solve these problems, the present invention provides an information processing method in which a pre-detection image, captured of the monitoring area before detection by a sensor monitoring the monitoring area, is input to a first generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generating AI or AI agent to generate pre-detection text information describing the context captured in the pre-detection image; a post-detection image, captured of the monitoring area after detection by the sensor, is input to a second generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the second generating AI or AI agent to generate post-detection text information describing the context captured in the post-detection image; information regarding the pre-detection text information and information regarding the post-detection text information are input to a third generating AI or AI agent that generates and outputs text information as a result of comparing a plurality of input text information, causing the third generating AI or AI agent to generate comparison result text information regarding the pre-detection text information and the post-detection text information, and outputting information regarding the comparison result text information. [Effects of the Invention]

[0015] The information processing system and information processing method according to the present invention enable the monitor to more accurately grasp the situation in the monitoring area. [Brief explanation of the drawing]

[0016] [Figure 1] This is a diagram showing the overall system configuration of Information Processing System 1. [Figure 2] This is a flowchart illustrating an example of how the learning process works. [Figure 3]It is a flowchart showing an example of the operation of the detection process. [Figure 4] (A) and (B) show an example of an image before detection, and FIGS. 4(C) and (D) show an example of an image after detection. [Figure 5] It is a flowchart showing an example of the operation of other learning processes. [Figure 6] It is a flowchart showing an example of the operation of other detection processes.The following describes the monitoring system according to the embodiment with reference to the drawings.

Mode for Carrying Out the Invention

[0017] Hereinafter, the monitoring system according to the embodiment will be described with reference to the drawings.

[0018] FIG. 1 is a diagram showing the overall system configuration of the information processing system 1. As shown in FIG. 1, the information processing system 1 includes a monitoring device 10, a center device 20, a server device 30, and the like. The number of each of the monitoring device 10, the center device 20, and the server device 30 is not limited to one, and may be plural. The monitoring device 10, the center device 20, and the server device 30 are communicably connected to each other via a network N. The network N is a wide area communication network such as an intranet or the Internet. The monitoring device 10 is installed in a monitoring target such as a house, a store, an office, a commercial facility, a factory, etc., and monitors one or a plurality of monitoring areas included in the monitoring target. The center device 20 is installed on a monitoring desk or the like in a monitoring center installed inside or outside the monitoring area, and aggregates and manages the monitoring results by the monitoring device 10. The server 30 is installed in a monitoring center or the like. The server 30 may be a cloud server in a wide area communication network such as the Internet.

[0019] The monitoring device 10 monitors abnormalities such as intrusion, fire, and fainting of people by devices such as sensors installed in the monitoring target. The monitoring device 10 includes an interface unit 11, an imaging unit 12, a monitoring sensor 13, a first input unit 14, a first output unit 15, a first communication unit 16, a first storage unit 17, and a first control unit 18.

[0020] The interface unit 11 has an interface circuit conforming to a serial bus standard such as USB, etc., and communicates with the imaging unit 12 and the monitoring sensor 13 to transmit and receive various signals. Note that the interface unit 11 may have an interface circuit conforming to a wired / wireless communication standard such as Ethernet (registered trademark), IEEE802.11, Bluetooth (registered trademark), etc., instead of the interface circuit conforming to the serial bus standard.

[0021] The imaging unit 12 is disposed in a monitoring area included in the monitoring target and images the monitoring area. The number of imaging units 12 in one monitoring area is not limited to one and may be plural. The imaging unit 12 includes a camera. The camera has, for example, a photoelectric conversion element such as a CCD element or a C-MOS element, an imaging optical system that forms an image on the photoelectric conversion element, and an A / D converter that amplifies an electrical signal output from the photoelectric conversion element and performs analog / digital (A / D) conversion. The imaging unit 12 sequentially generates input images at a predetermined frame period and outputs them to the first control unit 18.

[0022] The monitoring sensor 13 is placed in a monitoring area included in the monitored object and monitors the monitoring area to detect abnormalities such as intrusion by suspicious persons, fire, or collapse of a person. An abnormality in the monitoring area is an example of an event that the monitoring sensor 13 can detect. The number of monitoring sensors 13 in one monitoring area is not limited to one, and there may be multiple sensors. For example, the monitoring sensor 13 is a magnetic sensor that detects intrusion by detecting the operation of opening and closing parts such as doors or windows of a building. The monitoring sensor 13 may also be an infrared sensor that detects intrusion by suspicious persons, collapse of a person, etc., by detecting the movement or state of moving objects inside the monitoring area based on changes in the amount of infrared light received. The monitoring sensor 13 may also be an ultrasonic sensor that emits ultrasonic waves and detects the movement or state of moving objects inside the monitoring area based on changes in the magnitude of the received ultrasonic waves, etc., by detecting intrusion by suspicious persons, collapse of a person, etc. The monitoring sensor 13 may also be a temperature sensor that detects intrusion by suspicious persons, fire, collapse of a person, etc., by detecting the heat emitted by a human body or object. The monitoring sensor 13 may be an image sensor that detects the movement or state of moving objects within the monitoring area based on the difference signal between the captured image and a pre-stored background image, thereby detecting the intrusion of a suspicious person or a person collapsing. The monitoring sensor 13 may also be an acceleration sensor or vibration sensor installed in a wearable device worn by a person in the monitoring area, and may detect a person falling based on the output information of the acceleration sensor or vibration sensor. When an abnormality is detected in the monitoring area, the monitoring sensor 13 transmits a detection signal to the first control unit 18.

[0023] The first input unit 14 has an interface circuit that receives signals from an operating device such as a keyboard, mouse, or touch panel, and accepts operations from the user, and outputs a signal corresponding to the accepted operation to the first control unit 18.

[0024] The first output unit 15 has a display device such as a liquid crystal display or an organic EL display and an interface circuit that outputs images to the display device, and displays various information such as images and text according to instructions from the first control unit 18. The first output unit 15 also has an audio output device such as a speaker and an interface circuit that outputs audio to the audio output device, and outputs audio according to instructions from the first control unit 18.

[0025] The first communication unit 16 has a communication interface circuit conforming to wired / wireless communication standards such as Ethernet® and IEEE 802.11, and communicates with the center device 20 via the network N to send and receive various information.

[0026] The first storage unit 17 includes semiconductor memory, magnetic storage media, and / or optical storage media. The first storage unit 17 stores the code of the computer program executed by the first control unit 18 to control the operation of the monitoring device 10, as well as various data. The computer program is installed in the first storage unit 17 by known methods via a computer-readable storage medium such as a CD-ROM or DVD-ROM, or via a communication line. The computer program may also be distributed from a server and installed in the first storage unit 17. The first storage unit 17 also stores as data the shooting locations captured by the imaging unit 12, i.e., the attributes of each monitoring area (house, store, office, commercial facility, factory, etc.).

[0027] The first control unit 18 includes a processor such as a CPU or multiprocessor and its peripheral circuits, and the processor controls the operation of the monitoring device 10 by executing a computer program stored in the first storage unit 17. A DSP, LSI, ASIC, FPGA, etc. may be used as the first control unit 18.

[0028] The central device 20 aggregates and displays the monitoring results from the monitoring device 10, thereby notifying the person being monitored. The central device 20 includes a second input unit 21, a second output unit 22, a second communication unit 23, a second storage unit 24, and a second control unit 25.

[0029] The second input unit 21 has an interface circuit that receives signals from an operating device such as a keyboard, mouse, or touch panel, and accepts operations from the user, and outputs a signal corresponding to the accepted operation to the second control unit 25.

[0030] The second output unit 22 has a display device such as a liquid crystal display or an organic EL display and an interface circuit that outputs images to the display device, and displays various information such as images and text according to instructions from the second control unit 25. The second output unit 22 also has an audio output device such as a speaker and an interface circuit that outputs audio to the audio output device, and outputs audio according to instructions from the second control unit 25.

[0031] The second communication unit 23 has a communication interface circuit that conforms to wired / wireless communication standards such as Ethernet (registered trademark) and IEEE 802.11, and communicates with the monitoring device 10 and the server device 30 via the network N to send and receive various information.

[0032] The second storage unit 24 includes semiconductor memory, magnetic storage media, and / or optical storage media. The second storage unit 24 stores the code and various data of the computer program executed by the second control unit 25 to control the operation of the center device 20. The computer program is installed in the second storage unit 24 by known methods via a computer-readable storage medium such as a CD-ROM or DVD-ROM, or via a communication line. The computer program may also be distributed from a server or the like and installed in the second storage unit 24.

[0033] The second control unit 25 includes a processor such as a CPU or multiprocessor and its peripheral circuits. The processor controls the operation of the center device 20 by executing a computer program stored in the second storage unit 24. A DSP, LSI, ASIC, FPGA, or the like may be used as the second control unit 25. The second control unit 25 includes, as a functional module of a program running on the processor, an acquisition unit 251, a first generation unit 252, a second generation unit 253, a third generation unit 254, an output control unit 255, and a receiving unit 256, etc.

[0034] The server device 30 stores one or more types of generating AI or AI agents, generates text information using the generating AI or AI agents in accordance with a request from the center device 20, and transmits it to the monitoring device 10. The server device 30 includes a third input unit 31, a third output unit 32, a third communication unit 33, a third storage unit 34, and a third control unit 35.

[0035] The third input unit 31 has an interface circuit that receives signals from an operating device such as a keyboard, mouse, or touch panel, and accepts operations from the user, and outputs a signal corresponding to the accepted operation to the third control unit 35.

[0036] The third output unit 32 has a display device such as a liquid crystal display or an organic EL display and an interface circuit for outputting images to the display device, and displays various information such as images and text according to instructions from the third control unit 35. The third output unit 32 also has an audio output device such as a speaker and an interface circuit for outputting audio to the audio output device, and outputs audio according to instructions from the third control unit 35.

[0037] The third communication unit 33 has a communication interface circuit conforming to wired / wireless communication standards such as Ethernet (registered trademark) and IEEE 802.11, and communicates with the center device 20 via the network N to send and receive various information.

[0038] The third storage unit 34 includes semiconductor memory, magnetic storage media, and / or optical storage media. The third storage unit 34 stores the code and various data of the computer program executed by the third control unit 35 to control the operation of the server device 30. The computer program is installed in the third storage unit 34 by known methods via a computer-readable storage medium such as a CD-ROM or DVD-ROM, or via a communication line. The computer program may also be distributed from a server or the like and installed in the third storage unit 34.

[0039] The third memory unit 34 stores one or more artificial intelligence (AIs). Each artificial intelligence is pre-trained to generate and output information corresponding to the input information. The generative AI includes one or more VLMs (Vision Language Models) that, when given an image and natural language (text) as input, generate and output information corresponding to the input image and natural language. This generative AI generates and outputs text information that describes the context (events, situations, content, objects, etc.) depicted in the input image. Each generative AI, which is a VLM, used in each process described later may be the same model or different models. Examples of VLMs include LLaVA (Large Language and Vision Assistant), GPT (Generative Pre-trained Transformer)-4v, and GPT-4o. LLaVA is a VLM that combines Llama2, an LLM (Large Language Model), with CLIP, which includes a vision encoder and a text encoder. LLaVA outputs an answer that is appropriate to an image, given a single image and natural language indicating a question about that image. LLaVA converts images into feature vectors using CLIP's vision encoder, then converts these feature vectors into a format that can be input to Llama2, and also converts natural language into feature vectors using a tokenizer. LLaVA inputs the feature vectors converted from the image and the feature vectors converted from the natural language into Llama2 and outputs a response that matches the image.

[0040] Furthermore, the generative AI includes one or more LLMs (Large-Scale Language Models) that, when natural language is input, generate and output information corresponding to the input natural language. This generative AI generates and outputs text information as a result of comparing multiple input text pieces. Alternatively, this generative AI generates and outputs text information representing the common part of multiple input text pieces. Each generative AI, which is an LLM used in each process described later, may be the same model or different models. Examples of LLMs include Llama2, Vicuna, Mistral, and BERT (Bidirectional Encoder Representations from Transformers).

[0041] The third control unit 35 includes a processor such as a CPU or multiprocessor and its peripheral circuits. The processor controls the operation of the server device 30 by executing a computer program stored in the third storage unit 34. A DSP, LSI, ASIC, FPGA, or the like may be used as the third control unit 35.

[0042] Figure 2 is a flowchart illustrating an example of the operation of the learning process by the center device 20. The learning process is a process for memorizing (learning) the state in which no events (such as abnormalities) detected by the monitoring sensor 13 have occurred in the monitoring area. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program that has been pre-stored in the second storage unit 24.

[0043] First, the acquisition unit 251 acquires a pre-detection image of the monitoring area captured before detection by the monitoring sensor 13 (step S101). The pre-detection image is an input image of the monitoring area captured by the imaging unit 12 when the monitoring sensor 13 has not detected the event to be detected, that is, when the event to be detected by the monitoring sensor 13 has not occurred in the monitoring area. The pre-detection image may be a still image or a video. The acquisition unit 251 transmits a pre-detection image request signal to the monitoring device 10 via the second communication unit 23 to request the acquisition of the pre-detection image. When the first control unit 18 of the monitoring device 10 receives the pre-detection image request signal from the center device 20 via the first communication unit 16, it transmits the pre-detection image to the center device 20 via the first communication unit 16. The acquisition unit 251 acquires the pre-detection image by receiving it from the monitoring device 10 via the second communication unit 23. The first control unit 18 may also spontaneously transmit the pre-detection image to the center device 20 each time the imaging unit 12 generates a pre-detection image. Furthermore, the pre-detection image may be the image taken immediately before the sensor detected the abnormality. That is, the imaging unit 12 may always capture and store images for a certain period of time in the past, and when the monitoring sensor 13 detects an abnormality, it may transmit images of the few frames immediately preceding the detection to the center device 20 as the pre-detection image.

[0044] Next, the acquisition unit 251 uses known image processing techniques such as background subtraction or inter-frame subtraction to determine whether or not a change region exists in the acquired pre-detection image (step S102). If no change region exists in the pre-detection image, the acquisition unit 251 returns to step S101 and repeats the processing in steps S101 to S102.

[0045] On the other hand, if a change region exists in the pre-detection image, the first generation unit 252 obtains pre-detection text information related to the pre-detection image from the pre-detection image (step S103). For example, the first generation unit 252 obtains pre-detection text information related to the pre-detection image from the pre-detection image using a generation AI which is a VLM. This generation AI is an example of a first generation AI. The pre-detection text information is, for example, information that shows text (natural language) that describes the context shown in the pre-detection image. The first generation unit 252 generates pre-detection prompt information that includes instructions for generating pre-detection text information from the pre-detection image. The pre-detection prompt information is text information that includes instructions for describing the context of the pre-detection image. Preferably, the pre-detection prompt information includes conditions for describing the pre-detection image (conditions for specifying the text format to be output, conditions for points that require detailed explanation, conditions for points that should be ignored, etc.). The pre-detection prompt information may be, for example, "Please describe the situation shown in the image in detail using bullet points. Pay particular attention to the lighting conditions and the presence or absence of people."

[0046] The first generation unit 252 transmits a pre-detection request signal to the server device 30 via the second communication unit 23, instructing the generation AI to generate pre-detection text information by inputting the pre-detection image and pre-detection prompt information into the generation AI. The pre-detection request signal includes the pre-detection image and pre-detection prompt information. When the third control unit 35 of the server device 30 receives the pre-detection request signal from the center device 20 via the third communication unit 33, it inputs the pre-detection image and pre-detection prompt information included in the received pre-detection request signal into the generation AI, which is a VLM, causing the generation AI to generate pre-detection text information. The third control unit 35 transmits the generated pre-detection text information to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the pre-detection text information by receiving it from the server device 30 via the second communication unit 23. In this way, the first generation unit 252 inputs the pre-detection image into the generation AI and causes the generation AI to generate pre-detection text information related to the pre-detection image. As will be described later, the learning process is executed repeatedly, and the first generation unit 252 inputs multiple pre-detection images to the generation AI, causing the generation AI to generate multiple pre-detection text information.

[0047] Next, the first generation unit 252 reads normal text information from the second storage unit 24 (step S104). Normal text information is information generated / stored based on one or more pre-detection text information, and is stored in the second storage unit 24 in the processing described later. It is text information that indicates the context of the monitoring area when no abnormality has occurred. Normal text information is stored in groups for multiple scenes according to the context of the monitoring area. For example, normal text information representing the scene when the lights in the monitoring area are off, normal text information representing the scene when the lights in the monitoring area are on and a person is present, and normal text information representing the scene when the lights in the monitoring area are on and no person is present are stored separately. Normal text information representing each scene is generated / stored based on one or more pre-detection text information, as described later.

[0048] Next, the first generation unit 252 acquires differential text information that shows the difference between the acquired pre-detection text information and the normal text information for each scene read from the second storage unit 24 (step S105). The first generation unit 252 acquires differential text information using, for example, a generation AI which is an LLM. For each normal text information for each scene, the first generation unit 252 generates differential prompt information that includes an instruction to generate the difference between the pre-detection text information and the normal text information. The instruction includes the pre-detection text information and the normal text information. The differential prompt information is, for example, "Please describe the difference between [pre-detection text information] and [normal text information] in detail in bullet points." (The content shown in each text information is written in []).

[0049] The first generation unit 252 inputs the differential prompt information generated for each scene into the generation AI and sends a differential request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate differential text information for each scene. The differential request signal includes the differential prompt information for each scene. When the third control unit 35 of the server device 30 receives a differential request signal from the center device 20 via the third communication unit 33, it inputs the differential prompt information contained in the received differential request signal into the generation AI, which is an LLM, and causes the generation AI to generate differential text information. The third control unit 35 sends the generated differential text information for each scene to the center device 20 via the third communication unit 33. The first generation unit 252 acquires each differential text information by receiving it from the server device 30 via the second communication unit 23.

[0050] Next, the first generation unit 252 extracts the difference text information of the scene with the smallest difference between the pre-detection text information and the normal text information from the acquired difference text information of each scene as the minimum difference text information (step S106). The first generation unit 252 extracts the minimum difference text information using a generation AI, such as an LLM. The first generation unit 252 generates minimum difference prompt information that includes a command to extract the difference text information of the scene with the smallest difference between the pre-detection text information and the normal text information from the difference text information of each scene. The command includes the difference text information of each scene. The minimum difference prompt information is, for example, "Please select the one with the smallest difference from [Difference Text Information A], [Difference Text Information B]... However, please consider the presence or absence of people and lighting conditions as significant differences." (The content shown in each text information is written in []. Difference Text Information A indicates the difference text information of scene A, and Difference Text Information B indicates the difference text information of scene B.)

[0051] The first generation unit 252 sends a minimum difference request signal to the server device 30 via the second communication unit 23, instructing the generation AI to input minimum difference prompt information and extract minimum difference text information. The minimum difference request signal includes minimum difference prompt information. When the third control unit 35 of the server device 30 receives a minimum difference request signal from the center device 20 via the third communication unit 33, it inputs the minimum difference prompt information contained in the received minimum difference request signal to the generation AI, which is an LLM, causing the generation AI to extract minimum difference text information. The third control unit 35 transmits the extracted minimum difference text information to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the minimum difference text information by receiving it from the server device 30 via the second communication unit 23.

[0052] Next, the first generation unit 252 determines whether the extracted (acquired) minimum difference text information satisfies predetermined conditions (step S107). The predetermined conditions are, for example, that the items (differences) shown in the minimum difference text information are minor. The first generation unit 252 uses, for example, a generation AI which is LLM to determine whether the minimum difference text information satisfies predetermined conditions. The first generation unit 252 generates condition determination prompt information which includes an instruction to determine whether the minimum difference text information satisfies predetermined conditions. The instruction includes the minimum difference text information. The condition determination prompt information is, for example, "Is the difference shown in [minimum difference text information] minor? However, the presence or absence of people, the state of people, changes in lighting conditions, and dangerous conditions are not minor." (The content shown in each text information is described in []).

[0053] The first generation unit 252 inputs condition determination prompt information to the generation AI and sends a condition determination request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate a determination result. The condition determination request signal includes condition determination prompt information. When the third control unit 35 of the server device 30 receives a condition determination request signal from the center device 20 via the third communication unit 33, it inputs the condition determination prompt information included in the received condition determination request signal to the generation AI, which is an LLM, and causes the generation AI to generate a determination result. The third control unit 35 sends the generated determination result to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the determination result by receiving it from the server device 30 via the second communication unit 23. The first generation unit 252 determines whether the minimum difference text information satisfies the predetermined conditions based on whether the acquired determination result indicates that the minimum difference text information satisfies the predetermined conditions.

[0054] If the minimum difference text information does not meet the predetermined conditions, that is, if the difference between the pre-detection text information and each of the already stored normal text information is not negligible, the first generation unit 252 stores the pre-detection text information in the second storage unit 24 as normal text information representing a new scene (step S108). As a result, normal text information representing a new scene is added to the second storage unit 24 based on the newly acquired pre-detection image. Next, the first generation unit 252 returns to step S101 and repeats the processing from step S101 onwards.

[0055] On the other hand, if the minimum difference text information satisfies a predetermined condition, that is, if the difference between the pre-detection text information and the normal text information of any of the already stored scenes is minor, the first generation unit 252 considers that the pre-detection text information is information representing the context of that scene (hereinafter referred to as the "common scene"). The first generation unit 252 then extracts the common part between the pre-detection text information and the normal text information of the common scene (step S109). The first generation unit 252 extracts the common part using a generation AI, for example, an LLM. This generation AI is an example of a fourth generation AI. The first generation unit 252 generates common part extraction prompt information that includes an instruction to extract the common part between the pre-detection text information and the normal text information of the common scene. The instruction includes the pre-detection text information and the normal text information of the common scene. Common part extraction prompt information may include, for example, "Please extract the common part between [pre-detection text information] and [normal text information for the common scene]." (The content shown in each text information item is written within the brackets []).

[0056] The first generation unit 252 inputs common part extraction prompt information to the generation AI and sends a common part extraction request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate the common part between the pre-detection text information and the normal text information of the common scene. The common part extraction request signal includes common part extraction prompt information. When the third control unit 35 of the server device 30 receives the common part extraction request signal from the center device 20 via the third communication unit 33, it inputs the common part extraction prompt information included in the received common part extraction request signal to the generation AI, which is an LLM, and causes the generation AI to generate text information representing the common part between the pre-detection text information and the normal text information of the common scene. The third control unit 35 transmits the extracted text information representing the common part to the center device 20 via the third communication unit 33. The first generation unit 252 acquires the text information representing the common part between the pre-detection text information and the normal text information of the common scene by receiving it from the server device 30 via the second communication unit 23.

[0057] The first generation unit 252 inputs information about multiple pre-detection text information into the generation AI, causing the generation AI to generate text information representing the common part of the multiple pre-detection text information. This allows the information processing system 1 to extract the common part of the multiple pre-detection text information simply and with high accuracy. Alternatively, the first generation unit 252 may generate the common part by rule-based processing without having the generation AI generate it. For example, the first generation unit 252 may perform word or phrase-based matching comparisons on the multiple pre-detection text information and extract the matched text portion as the common part.

[0058] Next, the first generation unit 252 updates the normal text information of the common scene using the acquired text information indicating the common part (step S110). As a result, the content based on the newly acquired pre-detection image is reflected in the normal text information stored in the second storage unit 24.

[0059] In this way, the first generation unit 252 classifies the multiple pre-detection text information into multiple groups, each represented by normal text information that represents one or more scenes. This allows the information processing system 1 to compare the images of the monitoring area captured after detection by the monitoring sensor 13 with each pre-detection image for each classified group (scene) in the processing described later. Therefore, the information processing system 1 can determine with high accuracy whether or not the images of the monitoring area captured after detection by the monitoring sensor 13 have differences from the one or more pre-detection images corresponding to each scene.

[0060] Next, the first generation unit 252 returns to step S101 and repeats the processing from step S101 onward.

[0061] Figure 3 is a flowchart illustrating an example of the operation of the detection process by the center device 20. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program pre-stored in the second storage unit 24. The detection process is executed when the monitoring sensor 13 detects an abnormality or the like.

[0062] First, the acquisition unit 251 acquires a post-detection image of the monitoring area captured after detection by the monitoring sensor 13, and the imaging status of the monitoring area at the time the post-detection image was captured (step S201). The post-detection image is an input image of the monitoring area captured by the imaging unit 12 when the monitoring sensor 13 has detected an abnormality, that is, when an abnormality has occurred in the monitoring area. The post-detection image may be a still image or a video. The imaging status of the monitoring area is the shooting location and / or shooting time captured by the imaging unit 12. When the first control unit 18 of the monitoring device 10 receives an abnormality signal from the monitoring sensor 13, it transmits the input image received from the imaging unit 12 as a post-detection image, along with the shooting location or shooting time stored in the first storage unit 17, to the center device 20 via the first communication unit 16. The acquisition unit 251 acquires the post-detection image and the imaging status of the monitoring area by receiving them from the center device 20 via the second communication unit 23.

[0063] Next, the acquisition unit 251 uses known image processing techniques such as background subtraction or interframe subtraction to determine whether or not a change region exists in the acquired post-detection image (step S202). If no change region exists in the post-detection image, the acquisition unit 251 returns to step S201 and repeats the processing from steps S201 to S202.

[0064] On the other hand, if a region of change exists in the post-detection image, the second generation unit 253 obtains post-detection text information related to the post-detection image from the post-detection image (step S203). For example, the second generation unit 253 obtains post-detection text information related to the post-detection image from the post-detection image using a generation AI which is a VLM. This generation AI is an example of a second generation AI. The post-detection text information is, for example, information that shows text (natural language) that describes the context captured in the post-detection image. The second generation unit 253 generates post-detection prompt information that includes instructions for generating post-detection text information from the post-detection image. The post-detection prompt information is text information that includes instructions for describing the context of the post-detection image. Preferably, the post-detection prompt information includes conditions for describing the post-detection image. For example, the post-detection prompt information may be "Please describe the situation captured in the image in detail using bullet points. Pay particular attention to the lighting conditions and the presence or absence of people."

[0065] The second generation unit 253 transmits a post-detection request signal to the server device 30 via the second communication unit 23, instructing the generation AI to generate post-detection text information by inputting the post-detection image and post-detection prompt information into the generation AI. The post-detection request signal includes the post-detection image and post-detection prompt information. When the third control unit 35 of the server device 30 receives the post-detection request signal from the center device 20 via the third communication unit 33, it inputs the post-detection image and post-detection prompt information contained in the received post-detection request signal into the generation AI, which is a VLM, causing the generation AI to generate post-detection text information. The third control unit 35 transmits the generated post-detection text information to the center device 20 via the third communication unit 33. The second generation unit 253 acquires the post-detection text information by receiving it from the server device 30 via the second communication unit 23. In this way, the second generation unit 253 inputs the post-detection image into the generation AI and causes the generation AI to generate post-detection text information related to the post-detection image.

[0066] Next, the third generation unit 254 reads out each normal text information from the second storage unit 24 (step S204).

[0067] Next, the third generation unit 254 acquires differential text information that shows the difference between the acquired post-detection text information and the normal text information for each scene read from the second storage unit 24 (step S205). The third generation unit 254 acquires differential text information using a generation AI, for example, an LLM. This generation AI is an example of a third generation AI. For each normal text information of each scene, the third generation unit 254 generates differential prompt information that includes an instruction to generate the difference between the post-detection text information and the normal text information. The instruction includes the post-detection text information and the normal text information. The differential prompt information is, for example, "Please describe the difference between [post-detection text information] and [normal text information] in detail in bullet points." (The content shown in each text information is written in []).

[0068] The third generation unit 254 inputs the differential prompt information generated for each scene into the generation AI and sends a differential request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate differential text information for each scene. The differential request signal includes the differential prompt information for each scene. When the third control unit 35 of the server device 30 receives a differential request signal from the center device 20 via the third communication unit 33, it inputs the differential prompt information contained in the received differential request signal into the generation AI, which is an LLM, and causes the generation AI to generate differential text information. The third control unit 35 sends the generated differential text information for each scene to the center device 20 via the third communication unit 33. The third generation unit 254 acquires each differential text information by receiving it from the server device 30 via the second communication unit 23.

[0069] The content shown in the normal text information and the content shown in the pre-detection text information are examples of information related to the pre-detection text information, and the content shown in the post-detection text information are examples of information related to the post-detection text information. The differential text information is an example of comparison result text information related to the pre-detection text information and the post-detection text information. That is, the third generation unit 254 inputs information related to the pre-detection text information, information related to the post-detection text information, and information indicating instructions for generating comparison result text information related to the pre-detection text information and the post-detection text information to the generation AI, causing the generation AI to generate the comparison result text information. As a result, the information processing system 1 can easily and appropriately identify the cause of the detected anomaly from the image captured when the monitoring sensor 13 has detected an anomaly.

[0070] In particular, the third generation unit 254 generates information about pre-detection text information using multiple pre-detection text information, and causes the generation AI to generate text information that is a comparison result of the pre-detection text information and the post-detection text information. By using multiple pre-detection images and multiple pre-detection texts, the information processing system 1 can identify anomalies contained in images captured when the monitoring sensor 13 has detected an anomaly with high accuracy.

[0071] Furthermore, the third generation unit 254 generates information about pre-detection text information based on the common parts of multiple pre-detection text information, and causes the generation AI to generate text information resulting from a comparison between the pre-detection text information and the post-detection text information. By using the common parts of multiple pre-detection text information, a comparison can be made based on the common characteristics of the sensor before detection (normal state), and the information processing system 1 can identify abnormalities contained in images captured when the monitoring sensor 13 has detected an abnormality with higher accuracy.

[0072] Next, the third generation unit 254 extracts the difference text information of the scene with the smallest difference between the detected text information and the normal text information from the acquired difference text information of each scene as the minimum difference text information (step S206). The third generation unit 254 extracts the minimum difference text information using a generation AI, such as an LLM. The third generation unit 254 generates minimum difference prompt information that includes a command to extract the difference text information with the smallest difference between the detected text information and the normal text information from the difference text information of each scene. The command includes the difference text information of each scene. The minimum difference prompt information is, for example, "Please select the one with the smallest difference from [Difference Text Information A], [Difference Text Information B]... However, please consider the presence or absence of people and lighting conditions as significant differences." (The content shown in each text information is written in []. Difference Text Information A indicates the difference text information of scene A, and Difference Text Information B indicates the difference text information of scene B.)

[0073] The third generation unit 254 sends a minimum difference request signal to the server device 30 via the second communication unit 23, instructing the generation AI to input minimum difference prompt information and extract minimum difference text information. The minimum difference request signal includes minimum difference prompt information. When the third control unit 35 of the server device 30 receives a minimum difference request signal from the center device 20 via the third communication unit 33, it inputs the minimum difference prompt information contained in the received minimum difference request signal to the generation AI, which is an LLM, causing the generation AI to extract minimum difference text information. The third control unit 35 transmits the extracted minimum difference text information to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the minimum difference text information by receiving it from the server device 30 via the second communication unit 23.

[0074] The content shown in the pre-detection text information corresponding to the minimum difference text information is an example of information related to the pre-detection text information, and the content shown in the post-detection text information is an example of information related to the post-detection text information. Furthermore, the minimum difference text information is an example of comparison result text information related to the pre-detection text information and the post-detection text information. Specifically, the third generation unit 254 inputs information related to the pre-detection text information, information related to the post-detection text information, and information indicating instructions for generating comparison result text information related to the pre-detection text information and the post-detection text information to the generation AI, causing the generation AI to generate the comparison result text information. As a result, the information processing system 1 can appropriately identify the cause of the anomaly contained in the image captured when the monitoring sensor 13 has detected an anomaly.

[0075] Furthermore, the third generation unit 254 identifies the normal text information corresponding to the minimum difference text information based on the pre-detection text information included in the group (scene) that has the smallest difference from the post-detection text information, from among the groups (scenes) represented by each normal text information into which multiple pre-detection text information have been classified. By comparing this minimum difference text information with the post-detection text information, the third generation unit 254 can obtain a comparison result that focuses only on differences based on comparisons between the same scenes, without outputting differences based on comparisons between clearly different scenes (for example, differences in the scenes themselves such as "the brightness of the room is different") as comparison results, thereby enabling high-precision detection of anomalies in the post-detection text information.

[0076] Next, the third generation unit 254 determines whether or not an abnormality (an event detected by the monitoring sensor 13) has occurred in the monitoring area based on the extracted (acquired) minimum difference text information (step S207). The third generation unit 254 determines whether or not an abnormality has occurred in the monitoring area using, for example, a generation AI which is LLM. The third generation unit 254 generates abnormality determination prompt information which includes an instruction to determine whether or not the minimum difference text information indicates that an abnormality has occurred in the monitoring area. The instruction includes the minimum difference text information. The anomaly detection prompt information may include instructions to specify the reason why an anomaly was determined to have occurred in the monitoring area (the reason why the monitoring sensor 13 detected the target event) or why it was determined not to have occurred. In other words, the anomaly detection prompt information may include instructions to describe the reason why the minimum difference text information was determined to indicate an anomaly in the monitoring area or not. This allows the information processing system 1 to present information to the monitor so that they can correctly understand the status of the monitoring area. Furthermore, the anomaly detection prompt information may also include information regarding the imaging status of the monitoring area. This allows the information processing system 1 to determine with greater accuracy whether or not an anomaly has occurred in the monitoring area. If the imaging status includes information that identifies the location of the monitoring area and information that identifies the time of shooting, the information processing system 1 can take these circumstances into consideration to determine with greater accuracy whether or not an anomaly has occurred. For example, regarding a person lying face down on a desk, the information processing system 1 can determine that the person is taking a break if the shooting time is during a break, and that the person is unwell if it is outside of a break. Similarly, the information processing system 1 can determine that a person lying down is taking a break if the location is a bed, and that it has fallen if it is in a hallway. If the monitoring area is an office, the anomaly detection prompt information might be something like, "Does the difference shown in [Minimum Difference Text Information] indicate an anomaly or danger in the office? At a minimum, fire, medical emergencies, vandalism, etc., are considered danger signs, but not limited to those. Please also describe the reason why you determined that an anomaly occurred or did not occur in the office." (The content shown in each text information item is written in []).

[0077] The third generation unit 254 inputs abnormality determination prompt information to the generation AI and sends an abnormality determination request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate a determination result. The abnormality determination request signal includes abnormality determination prompt information. When the third control unit 35 of the server device 30 receives an abnormality determination request signal from the center device 20 via the third communication unit 33, it inputs the abnormality determination prompt information contained in the received abnormality determination request signal to the generation AI, which is an LLM, and causes the generation AI to generate a determination result. The third control unit 35 sends the generated determination result to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the determination result by receiving it from the server device 30 via the second communication unit 23. The third generation unit 254 determines whether or not an abnormality has occurred in the monitoring area based on whether or not the acquired determination result indicates that an abnormality has occurred in the monitoring area based on whether or not the minimum difference text information indicates that an abnormality has occurred in the monitoring area.

[0078] In this way, the third generation unit 254 inputs information regarding pre-detection text information, information regarding post-detection text information, and information indicating instructions for generating information indicating whether or not an event targeted by the monitoring sensor 13 has occurred in the monitoring area and / or the reason why the monitoring sensor 13 detected the event, causing the generation unit to generate information indicating whether or not an event targeted by the monitoring sensor 13 has occurred in the monitoring area and / or the reason why the monitoring sensor 13 detected the event. As a result, the information processing system 1 can appropriately determine whether or not an abnormality has occurred in the monitoring area or present the reason for determining that an abnormality has occurred in the monitoring area to the monitor.

[0079] Furthermore, the third generation unit 254 inputs information regarding pre-detection text information, information regarding post-detection text information, information regarding the imaging status of the monitoring area, and information indicating instructions for generating information indicating whether or not an event targeted by the monitoring sensor 13 has occurred in the monitoring area or the reason why the monitoring sensor 13 detected the event, to the generation AI. Considering the imaging status of the monitoring area, the generation unit causes the generation AI to generate information indicating whether or not an event targeted by the monitoring sensor 13 has occurred in the monitoring area or the reason why the monitoring sensor 13 detected the event. As a result, the information processing system 1 can appropriately determine whether or not an abnormality has occurred in the monitoring area or present the reason for determining that an abnormality has occurred in the monitoring area to the monitor.

[0080] If an abnormality occurs in the monitoring area, the output control unit 255 outputs the acquired determination result, i.e., that an abnormality has occurred in the monitoring area and / or the reason for determining that an abnormality has occurred, by displaying or outputting it as sound on the second output unit 22 (step S208). The output control unit 255 may also output the statement that an abnormality has occurred in the monitoring area and / or the reason for determining that an abnormality has occurred, with predetermined information added or subtracted. This allows the monitor to appropriately identify whether or not an abnormality has occurred in the monitoring area and / or the reason for determining that an abnormality has occurred, and the information processing system 1 can monitor the monitoring area more reliably.

[0081] The fact that an anomaly (an event detected by the monitoring sensor 13) has occurred in the monitoring area, and / or the reason why it was determined that an anomaly has occurred in the monitoring area (the reason why the monitoring sensor 13 detected the event), is an example of information related to the comparison result text information. The output control unit 255 may also output the comparison result text information itself. This allows the monitor to appropriately identify the type of anomaly that has occurred, and the information processing system 1 can monitor the monitoring area more reliably.

[0082] Next, the third generation unit 254 returns to step S201 and repeats the processing from step S201 onward.

[0083] On the other hand, if no abnormality occurs in the monitoring area, the third generation unit 254 extracts the common portion between the post-detection text information and the normal text information corresponding to the extracted minimum difference text information (step S209). The third generation unit 254 extracts the common portion using a generation AI, for example, an LLM. The third generation unit 254 generates common portion extraction prompt information that includes an instruction to extract the common portion between the post-detection text information and the normal text information corresponding to the minimum difference text information. The instruction includes the post-detection text information and the normal text information corresponding to the minimum difference text information. The common portion extraction prompt information is, for example, "Please extract the common portion of [post-detection text information] and [normal text information]." (The content shown in each text information is written in []).

[0084] The third generation unit 254 inputs common part extraction prompt information to the generation AI and sends a common part extraction request signal to the server device 30 via the second communication unit 23, causing the generation AI to extract the common part between the pre-detection text information and the normal text information corresponding to the minimum difference text information. The common part extraction request signal includes common part extraction prompt information. When the third control unit 35 of the server device 30 receives the common part extraction request signal from the center device 20 via the third communication unit 33, it inputs the common part extraction prompt information included in the received common part extraction request signal to the generation AI, which is an LLM, causing the generation AI to extract the common part between the post-detection text information and the normal text information corresponding to the minimum difference text information. The third control unit 35 transmits the extracted common part to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the common part between the post-detection text information and the normal text information corresponding to the minimum difference text information by receiving it from the server device 30 via the second communication unit 23.

[0085] Next, the third generation unit 254 updates the normal text information of the scene corresponding to the minimum difference text information with text information indicating the acquired common part (step S210). As a result, the normal text information stored in the second storage unit 24 is updated with content based on the post-detection image newly acquired when no abnormality has occurred in the monitoring area. Next, the third generation unit 254 returns to step S201 and repeats the processing from step S201 onwards.

[0086] Note that the process in step S102 in Figure 2 and / or the process in step S202 in Figure 3 may be omitted. Also, in the process in step S102 in Figure 2 and / or the process in step S202 in Figure 3, the acquisition unit 251 may use known image processing techniques to detect a predetermined object such as a person or a car from the pre-detection image or the post-detection image, and proceed to step S103 or S203 only if the predetermined object is detected.

[0087] Furthermore, if it is determined in step S207 of Figure 3 that no abnormality has occurred in the monitoring area, the output control unit 255 may also output the determination result, that is, that no abnormality has occurred in the monitoring area, and / or the reason for determining that no abnormality has occurred in the monitoring area.

[0088] Furthermore, the process in step S109 in Figure 2 and / or the process in step S209 in Figure 3 may be omitted, and the normal text information may be updated with the latest pre-detection text information or post-detection text information. Also, the processes in steps S109 to S110 in Figure 2 and / or the processes in steps S209 to S210 in Figure 3 may be omitted, and the normal text information may not be updated. Also, the processes in steps S104 to S110 in Figure 2 and / or the processes in steps S204 to S206 and S209 to S210 in Figure 3 may be omitted. In that case, the third generation unit 254 acquires differential text information showing the difference between the post-detection text information and the pre-detection text information acquired immediately before, in the same manner as the process in step S205, and determines whether or not an abnormality has occurred in the monitoring area based on the acquired differential text information, in the same manner as the process in step S207.

[0089] Furthermore, the learning process shown in Figure 2 may be run continuously. In that case, the normal text information will continue to be updated as long as the monitoring sensor 13 does not detect an abnormality. Therefore, when the monitoring sensor 13 detects an abnormality and the detection process shown in Figure 3 is executed, the normal text information is updated with the pre-detection image of the monitoring area taken immediately before detection by the monitoring sensor 13. That is, in the learning process, the first generation unit 252 inputs one or more pre-detection images of the monitoring area taken immediately before detection by the monitoring sensor 13 to the generation AI, causing the generation AI to generate pre-detection text information related to the pre-detection image. In this case, the information processing system 1 can compare the image taken when an abnormality occurs in the monitoring area with the image taken immediately before the abnormality occurred in the monitoring area, and can identify changes based on the abnormality that occurred in the monitoring area with higher accuracy.

[0090] Furthermore, the prompt information used in the learning and detection processes may not be pre-set but may be specified by the monitor. In this case, the reception unit 256 accepts the specification of each prompt information entered by the monitor using the second input unit 21 at any time. This allows the monitor to flexibly change each prompt information according to the status of the monitoring area, and the information processing system 1 can improve the convenience of the monitor and improve the security of the monitoring area.

[0091] Figures 4(A) and (B) show examples of images before detection, and Figures 4(C) and (D) show examples of images after detection.

[0092] The pre-detection images shown in Figures 4(A) and (B), and the post-detection images shown in Figures 4(C) and (D), are images of the office taken at different times. In the pre-detection images in Figures 4(A) and (B), and in the post-detection image in Figure 4(C), person P is working, but in the post-detection image in Figure 4(D), person P is lying down.

[0093] The pre-detection text information generated from the pre-detection image in Figure 4(A) describes the scene as follows: "Office setting: There is one person. He is wearing a white shirt and blue pants and is working at a computer. There is a computer on the desk. Lighting: It appears to be typical office lighting, and it is presumed that no special lighting is being used. Fluorescent lights and incandescent bulbs installed on the ceiling are likely being used. Floor: It appears to be a normal office floor, plain and simple in design." On the other hand, the pre-detection text information generated from the pre-detection image in Figure 4(B) describes the scene as follows: "An office scene is depicted. In the center is a desk with three black monitors and a laptop computer. Person: One person is sitting and working with deep concentration (using a laptop computer). Lighting: The lighting setup is unknown. Wall: White wallpaper." The normal text information, which is a common part extracted from each pre-detection text information, describes the scene as follows: "There is a computer on the desk. One person is working (using a laptop computer)." As shown above, the pre-detection images in Figures 4(A) and 4(B) are similar to each other, but the content described in the pre-detection text information generated from each pre-detection image is slightly different. By extracting common parts from each pre-detection image, the information processing system 1 can appropriately extract the characteristics of the monitoring area when no abnormality has occurred.

[0094] The post-detection text information generated from the post-detection image in Figure 4(C) states: "Situation: A man is sitting at a desk and working on a computer. Presence of people: One man is visible. Three computer screens are present, and a laptop is open in front of the man. Lighting: The entire room relies on natural light. Since no windows are visible, it is unknown whether there is external light." The difference text information, which is the difference between the above normal text information and this post-detection text information, states: "The laptop is open in front of the man. Lighting: The entire room relies on natural light." The result of the determination of whether or not an abnormality has occurred in the monitoring area states: "No abnormality. Reason: The laptop being open in front of the man is a common sight during work and is not unusual. The entire room relies on natural light, which is a normal office lighting environment and is not unusual." Based on these determination results, the monitoring officer can easily understand that the monitoring sensor 13 had made a false positive.

[0095] The post-detection text information generated from the post-detection image in Figure 4(D) states: "Location: Office. Presence of people: 1 person (male). Condition: The man is lying on the floor. Clothing: Wearing a white shirt and black pants. Surrounding environment: The man is located near a desk, with a computer screen and laptop on the desk. Lighting conditions: The office has standard lighting settings, with no particularly noticeable lighting fixtures or special lighting effects." The difference text information, which is the difference between the above normal text information and this post-detection text information, states: "A person is lying on the floor." The result of the determination of whether or not an abnormality has occurred in the monitoring area states: "Abnormality detected. Reason: It is unnatural for a person to be lying on the floor in a normal office." Based on these determination results, the monitoring officer can easily understand that the cause detected by the monitoring sensor 13 was a person collapsing.

[0096] In this way, if there is a difference between the pre-detection image and the post-detection image, the information processing system 1 uses the generated AI to notify the monitor whether or not an anomaly has occurred in the monitoring area and / or the reason for that determination. As a result, the monitor can easily and accurately grasp the situation in the monitoring area, regardless of their skills, experience, or prior knowledge of the monitoring area, and can respond appropriately to any anomalies occurring in the monitoring area.

[0097] As explained above, the information processing system 1 generates and outputs text information resulting from a comparison between pre-detection text information generated by the generating AI from the pre-detection image and post-detection text information generated by the generating AI from the post-detection image. The information processing system 1 can appropriately determine the difference between the pre-detection image and the post-detection image by utilizing the knowledge that the generating AI has learned. Therefore, with the information processing system 1, the monitor can more accurately grasp the situation of the monitoring area. Furthermore, by comparing with normal text information, the information processing system 1 can compare with one or more scenes that represent the normal state of the monitoring area, and the monitor can more accurately grasp the situation of the monitoring area while also considering the difference from the normal state predetermined for the monitoring area.

[0098] Figures 5 and 6 are flowcharts illustrating other embodiments.

[0099] Figure 5 is a flowchart showing an example of the operation of the learning process according to another embodiment. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program that is pre-stored in the second storage unit 24. Since the processing in steps S301 to S302 in Figure 5 is the same as the processing in steps S101 to S102 in Figure 2, only the processing in steps S303 to S304 will be described below.

[0100] If a change region exists in the pre-detection image in step S302, the first generation unit 252 acquires pre-detection feature information related to the pre-detection image from the pre-detection image (step S303). In the same manner as in step S103, the first generation unit 252 transmits a pre-detection request signal to the server device 30 via the second communication unit 23, and the third control unit 35 of the server device 30 causes the generation AI to generate pre-detection text information. However, the third control unit 35 acquires the feature quantities of the pre-detection image, which are intermediate products generated by the generation AI, as pre-detection feature information and transmits them to the center device 20 via the third communication unit 33. The intermediate products are obtained from the feature quantities output from the vision encoder of the VLM, which is the generation AI that generates pre-detection text information. The third control unit 35 acquires the feature quantities that the text decoder of the VLM can recognize as pre-detection feature information. The first generation unit 252 acquires the pre-detection feature information by receiving it from the server device 30 via the second communication unit 23. In this way, when the first generation unit 252 inputs the pre-detection image to the generation AI, it generates feature information generated by the generation AI.

[0101] Next, the first generation unit 252 stores the acquired pre-detection feature information as normal feature information in the second storage unit 24 (step S304). Then, the first generation unit 252 returns to step S301 and repeats the process from step S301 onwards. As this process is repeated, normal feature information consisting of multiple pre-detection feature information is stored in the second storage unit 24.

[0102] Figure 6 is a flowchart showing an example of the operation of the detection process according to another embodiment. This flowchart is executed mainly by the second control unit 25 in cooperation with each element of the center device 20, based on a program that is pre-stored in the second storage unit 24. Since the processing of steps S401 to S402 in Figure 6 is the same as the processing of steps S201 to S202 in Figure 3, only the processing of steps S403 to S410 will be described below.

[0103] If a change region exists in the post-detection image in step S402, the second generation unit 253 acquires post-detection feature information related to the post-detection image from the post-detection image (step S403). In the same manner as in step S203, the second generation unit 253 transmits a post-detection request signal to the server device 30 via the second communication unit 23, and the third control unit 35 of the server device 30 causes the generation AI to generate post-detection text information. However, the third control unit 35 acquires the feature quantities of the post-detection image, which are intermediate products generated by the generation AI, as post-detection feature information and transmits them to the center device 20 via the third communication unit 33. The third control unit 35 acquires the feature quantities that can be recognized by the text decoder of the generation AI, VLM, as post-detection feature information. The second generation unit 253 acquires the post-detection feature information by receiving it from the server device 30 via the second communication unit 23.

[0104] Next, the third generation unit 254 classifies the normal feature information, which consists of multiple pre-detection feature information stored in the second storage unit 24, into multiple groups by clustering (step S404). The third generation unit 254 clusters each normal feature information using a known statistical method such as the k-means method. Since the features are classified more flexibly and appropriately than text, the information processing system 1 can appropriately classify the features of images of monitoring areas where no abnormalities have occurred, according to environmental conditions.

[0105] Next, the third generation unit 254 identifies the group of normal feature information that best approximates the post-detection feature information obtained in step S403 (step S405). For example, the third generation unit 254 calculates the centroid position of each group in the feature space and identifies the group having the centroid position with the smallest Euclidean distance from the post-detection feature information as the group that best approximates the post-detection feature information.

[0106] Next, the third generation unit 254 acquires the closest-to-the-nearest text information corresponding to the identified group and the post-detection text information corresponding to the post-detection feature information (step S406). For example, the third generation unit 254 calculates the feature corresponding to the centroid position of the identified group as the closest-to-the-nearest feature information. The third generation unit 254 sends a decode request signal to the server device 30 via the second communication unit 23 to cause the generation AI to generate (decode) the text information corresponding to the closest-to-the-nearest feature information and the text information corresponding to the post-detection feature information. The decode request signal includes the closest-to-the-nearest feature information and the post-detection feature information. When the third control unit 35 of the server device 30 receives the decode request signal from the center device 20 via the third communication unit 33, it inputs the closest-to-the-nearest feature information contained in the received decode request signal to the generation AI (VLM text decoder), which is an LLM, and causes the generation AI to generate the closest-to-the-nearest text information corresponding to the closest-to-the-nearest feature information. Furthermore, the third control unit 35 inputs the post-detection feature information contained in the received decode request signal to the generation AI (VLM text decoder), which is an LLM, and causes the generation AI to generate post-detection text information corresponding to the post-detection feature information. The third control unit 35 transmits the generated similar text information and post-detection text information to the center device 20 via the third communication unit 33. The third generation unit 254 acquires the similar text information and post-detection text information by receiving them from the server device 30 via the second communication unit 23.

[0107] Next, the third generation unit 254 determines whether or not an abnormality has occurred in the monitoring area based on the acquired similar text information and post-detection text information (step S407). The third generation unit 254 determines whether or not an abnormality has occurred in the monitoring area using a generation AI, for example, LLM. The third generation unit 254 generates abnormality determination prompt information and obtains the determination result in the same manner as the process in step S207 of Figure 3. However, the abnormality determination prompt information in this embodiment includes an instruction to determine whether or not the post-detection text information indicates that an abnormality has occurred in the monitoring area. The instruction includes similar text information and post-detection text information. The instruction may also include information regarding the imaging status. If the monitoring area is an office, the anomaly detection prompt information may be something like: "Does the difference shown in [Recently Similar Text Information] and [Target Text Information] indicate an anomaly or danger in the office? At a minimum, fire, sudden illness, vandalism, etc., should be considered danger signs, but not limited to these. Pay particular attention to the presence and condition of people when making your determination. Please also describe the reason why you determined that an anomaly occurred in the office or why you determined that no anomaly occurred." (The content shown in each text information is written in []).

[0108] Nearest similar text information is an example of information related to pre-detection text information. That is, when the third generation unit 254 inputs a pre-detection image to the generation AI, it generates information related to pre-detection text information based on the feature information generated by the generation AI. In this case as well, the information processing system 1 can appropriately determine whether or not an abnormality has occurred in the monitoring area.

[0109] If an abnormality occurs in the monitoring area, the output control unit 255 outputs the acquired judgment result in the same manner as the process in step S208 (step S408). Next, the third generation unit 254 returns to step S401 and repeats the process from step S401 onwards.

[0110] On the other hand, if no abnormality occurs in the monitoring area, the third generation unit 254 stores the post-detection feature information as normal feature information in the second storage unit 24 (step S409). Next, the third generation unit 254 returns to step S401 and repeats the processing from step S401 onwards.

[0111] As explained above, even when the information processing system 1 uses feature information generated by the generating AI, the monitor can more accurately grasp the situation of the monitoring area. In the other embodiments described above, the third generation unit 254 grouped the normal feature information in the detection process, but instead, the normal feature information may be grouped in the learning process. That is, the third generation unit 254 may group the normal features for each scene by performing a clustering process similar to the process in step S404 at the end of the learning process, and then extract the closest similarity group from each group (scene) in step S405 of the detection process.

[0112] Although preferred embodiments have been described above, the embodiments are not limited to the examples described above. For example, the information processing system 1 may have an AI agent generate each piece of information instead of a generating AI. In that case, the third storage unit 34 stores one or more AI agents. Each AI agent is pre-trained to generate and output information corresponding to the input information.

[0113] The AI ​​agent includes one or more autonomous AI agents using VLM that generate and output information corresponding to the input image and natural language when an image and natural language are input. This AI agent generates and outputs text information that describes the context depicted in the input image. The AI ​​agent also includes one or more autonomous AI agents using LLM that generate and output information corresponding to the input natural language when natural language is input. This AI agent generates and outputs text information as a result of comparing multiple input text information. Alternatively, this AI agent generates and outputs text information representing the common part of multiple input text information. When given a goal, the autonomous AI agent outputs information that achieves the goal by having a generating AI generate tasks to achieve that goal, collecting information to allow the generating AI to execute the generated tasks, and repeating the process of having the generating AI execute the tasks. The AI ​​agent may also autonomously acquire environmental information such as temperature and lighting of the monitored object (monitoring area), and output information that achieves the goal while considering the acquired environmental information.

[0114] In each process where a VLM-based generative AI is used, as shown in Figures 2, 3, 5, and 6, an AI agent using VLM is used, and in each process where an LLM-based generative AI is used, an AI agent using LLM is used. Each AI agent used in place of the 1st to 4th generative AIs is an example of the 1st to 4th AI agents, respectively. In each process, no prompt information is generated; instead, an AI agent generated to achieve the goal corresponding to each prompt information is used. Each VLM-based AI agent used in each process may be the same agent or different agents. Each LLM-based AI agent used in each process may be the same agent or different agents.

[0115] Even when the information processing system 1 uses an AI agent, the monitor can more accurately grasp the situation in the monitoring area.

[0116] Furthermore, in the information processing system 1, the learning process and detection process may be performed by the monitoring device 10 instead of the center device 20. In that case, the first control unit 18 of the monitoring device 10 has the parts of the second control unit 25 and performs the learning process and detection process. Alternatively, the learning process and detection process may be performed by the server device 30 instead of the center device 20. In that case, the third control unit 35 of the server device 30 has the parts of the second control unit 25 and performs the learning process and detection process. In that case, in step S208 in Figure 3 and step S408 in Figure 6, the server device 30 may transmit the determination result to the center device 20 or the monitoring device 10 via the third communication unit 33 and output it to the second output unit 22 or the first output unit 15.

[0117] Furthermore, each generating AI and / or each AI agent may be stored in the storage unit of the device that performs the learning and detection processing, rather than in the third storage unit 34 of the server device 30. In that case, when the device that performs the learning and detection processing uses each generating AI and / or each AI agent, it does not send each request signal to the server device 30, but instead causes each generating AI and / or each AI agent stored in the storage unit of the device to generate the information.

[0118] Furthermore, in the above embodiment, the first generation unit 252 generates a common portion based on multiple pre-detection text information it has generated, and the third generation unit 254 generates comparison result text information based on that common portion. However, the system is not limited to this, and the third generation unit 254 may generate a common portion using multiple pre-detection text information generated by the first generation unit 252, and then generate comparison result text information. In this case, the third generation unit 254 may generate the common portion using a fourth generation AI or AI agent, similar to the first generation unit 252. Alternatively, the third generation unit 254 may generate the common portion without using a fourth generation AI or AI agent. [Explanation of Symbols]

[0119] 1 Information processing system, 10 Monitoring device, 20 Center device, 30 Server device, 251 Acquisition unit, 252 First generation unit, 253 Second generation unit, 254 Third generation unit, 255 Output control unit, 256 Reception unit

Claims

1. A first generation unit inputs a pre-detection image of the monitoring area, which is captured before detection by a sensor monitoring the monitoring area, to a first generation AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generation AI or AI agent to generate pre-detection text information describing the context captured in the pre-detection image. A second generation unit inputs the post-detection image, which captures the monitoring area after detection by the sensor, to a second generation AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the second generation AI or AI agent to generate post-detection text information describing the context captured in the post-detection image. The third generation unit inputs the information regarding the pre-detection text information and the information regarding the post-detection text information to a third generation AI or AI agent that generates and outputs text information as a result of comparing multiple input text information, causing the third generation AI or AI agent to generate the comparison result text information regarding the pre-detection text information and the post-detection text information. An output unit that outputs information related to the comparison result text information, An information processing system characterized by having the following features.

2. The third generation unit inputs the information regarding the pre-detection text information and the information regarding the post-detection text information to the third generation AI or AI agent, causing the third generation AI or AI agent to generate information indicating whether or not the event targeted by the sensor has occurred in the monitoring area or the reason why the sensor detected the event. The information processing system according to claim 1, wherein the output unit outputs information regarding whether or not the event has occurred in the monitoring area or the reason for the detection.

3. The third generation unit inputs information regarding the pre-detection text information, information regarding the post-detection text information, and information regarding the imaging status of the monitoring area to the third generation AI or AI agent, and causes the third generation AI or AI agent to generate information indicating whether or not an event targeted by the sensor has occurred in the monitoring area, or the reason why the sensor detected the event, taking the imaging status into consideration. The information processing system according to claim 1, wherein the output unit outputs information regarding whether or not the event has occurred in the monitoring area or the reason for the detection.

4. The aforementioned pre-detection image is an image taken when no event to be detected by the sensor has occurred in the monitoring area. The first generation unit inputs a plurality of pre-detection images to the first generation AI or AI agent, and causes the first generation AI or AI agent to generate a plurality of pre-detection text information. The information processing system according to claim 1, wherein the third generation unit causes the third generation AI or AI agent to generate the comparison result text information using the plurality of pre-detection text information.

5. The information processing system according to claim 4, wherein the third generation unit generates a common portion of the multiple pre-detection text information using the multiple pre-detection text information, and causes the third generation AI or AI agent to generate the comparison result text information based on the common portion.

6. The first generation unit inputs a plurality of pre-detection images to the first generation AI or AI agent, causes the first generation AI or AI agent to generate a plurality of pre-detection text information, inputs the plurality of pre-detection text information to a fourth generation AI or AI agent that generates and outputs text information representing the common part of the plurality of input text information, causes the fourth generation AI or AI agent to generate text information representing the common part of the plurality of pre-detection text information, The information processing system according to claim 1, wherein the third generation unit causes the third generation AI or AI agent to generate the comparison result text information based on the common part.

7. When the first generation unit inputs the pre-detection image to the first generation AI or AI agent, it generates feature information generated by the first generation AI or AI agent. The information processing system according to any one of claims 1 to 3, wherein the third generation unit generates information relating to the pre-detection text information based on the feature information.

8. The pre-detection image of the monitoring area, captured before detection by the sensor monitoring the monitoring area, is input to a first generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the first generating AI or AI agent to generate pre-detection text information describing the context captured in the pre-detection image. The post-detection image, captured after the sensor's detection of the monitoring area, is input to a second generating AI or AI agent that generates and outputs text information describing the context captured in the input image, causing the second generating AI or AI agent to generate post-detection text information describing the context captured in the post-detection image. The information regarding the pre-detection text information and the information regarding the post-detection text information are input to a third generating AI or AI agent that generates and outputs text information as a result of comparing multiple input text information, causing the third generating AI or AI agent to generate text information as a result of comparing the information regarding the pre-detection text information and the information regarding the post-detection text information. Outputs information related to the comparison result text information. An information processing method characterized by the following:

Citation Information

Patent Citations

  • Monitoring device

    JP2006338187A