Information processing device, method for controlling the information processing device, and control program for the information processing device
A generative AI model compares time-series images from surveillance cameras to detect anomalies, improving the detection of small objects and reducing guard workload by generating text-based notifications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOFTBANK CORPORATION
- Filing Date
- 2025-06-17
- Publication Date
- 2026-04-16
AI Technical Summary
Existing security systems using surveillance cameras struggle to detect stationary objects, particularly those below a certain size, and are inconsistent in identifying abnormalities, leading to inefficiencies in building management and increased workload for security guards.
Utilizing a generative AI model to compare time-series images from surveillance cameras, detecting differences and generating notifications in text form about anomalies, allowing security guards to focus on high-priority events.
Enhances the detection of small, stationary objects and reduces the workload of security guards by providing accurate, text-based notifications of anomalies, enabling rapid response to critical events.
Smart Images

Figure 0007847260000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing apparatus, a control method for an information processing apparatus, and a control program for an information processing apparatus. [Background Art]
[0002] Conventionally, in a security system using a surveillance camera, there has been disclosed a surveillance device that acquires an image captured by the surveillance camera and generates a surveillance result according to the image captured by the surveillance camera using learned surveillance logic (for example, Patent Document 1). [Prior Art Documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2022-055229 [Summary of the Invention] [Means for Solving the Problems]
[0004] An information processing apparatus according to an embodiment of the present invention includes an image acquisition unit that acquires an image of a security target imaged by an imaging device, and a prompt including an instruction to output, in text, a result of comparing a reference image preset as a reference image among the images of the security target, a latest image that is the latest image in time series, and a past image that is an image acquired earlier than the latest image to a generation AI model, an instruction unit that inputs the prompt to the generation AI model, and a notification unit that generates a notification based on the output of the generation AI model.
[0005] In the information processing apparatus according to an embodiment of the present invention, in the instruction, the instruction unit may use, as the past image, an image located immediately before the latest image in time series.
[0006] In an information processing device according to one embodiment of the present invention, the image acquisition unit may acquire an image of the security target from the imaging device when a predetermined condition is met, and the instruction unit may input a prompt to the generation AI model when a new image of the security target is acquired.
[0007] In an information processing device according to one embodiment of the present invention, the image acquisition unit may acquire an image of the security target from the imaging device in accordance with predetermined conditions, such as the elapsed time since the previous image was acquired.
[0008] In an information processing device according to one embodiment of the present invention, the instruction unit may input a prompt to the generating AI model that includes an instruction to output the type of difference between the reference image, the latest image, and the past image based on the result of comparing the reference image, the latest image, and the past image.
[0009] In an information processing device according to one embodiment of the present invention, if multiple differences are detected, the instruction unit may input a prompt to the generating AI model that includes an instruction to enumerate the types of differences in order of priority set according to the type of difference.
[0010] In an information processing device according to one embodiment of the present invention, the notification unit may output a notification to the communication terminal of the user monitoring the security target.
[0011] An information processing device according to one embodiment of the present invention may further include a receiving unit that receives the selection of an image to be set as a reference image from among the images of the object to be guarded.
[0012] In an information processing device according to one embodiment of the present invention, the instruction unit may issue an instruction to output the result of comparing a reference image, the latest image, and a past image, and may input the reference image, the latest image, and the past image to the generated AI model at different timings.
[0013] In an information processing device according to one embodiment of the present invention, the notification unit may generate a notification that includes the latest image.
[0014] A control method for an information processing device according to one embodiment of the present invention includes the steps of: the information processing device acquiring an image of a security target captured by an imaging device; inputting a prompt to a generating AI model that includes an instruction to output in text the result of comparing a reference image, which is set in advance as a reference image among the images of the security target, the latest image which is the most recent image in the time series, and past images which are images acquired in the time series earlier than the latest image; and generating a notification based on the output of the generating AI model.
[0015] A control program for an information processing device according to one embodiment of the present invention provides the information processing device with the following functions: a function to acquire an image of a security target captured by an imaging device; a function to input a prompt to a generating AI model that includes an instruction to output in text the result of comparing an image of the security target with a pre-set reference image, the latest image which is the most recent image in the time series, and past images which are images acquired earlier in the time series than the latest image; and a function to generate a notification based on the output of the generating AI model. [Brief explanation of the drawing]
[0016] [Figure 1] Figure 1 is a schematic diagram of the security support system configuration according to one embodiment of the present invention. [Figure 2] Figure 2 shows an example of a sequence between a server, an imaging device, a generation AI system, and a communication terminal in one embodiment of the present invention. [Figure 3] Figures 3(a) and 3(b) are schematic diagrams illustrating one embodiment of the present invention. [Figure 4] Figure 4 is an example of a functional block diagram of a server (information processing device) according to one embodiment of the present invention. [Figure 5] Figure 5 shows an example of the control flow of a server according to one embodiment of the present invention. [Modes for carrying out the invention]
[0017] Hereafter, an embodiment of the invention described herein (also referred to as the present invention) will be explained using the figures. Note that the figures are examples only, and the present invention is not limited to those shown in the figures. For example, the illustrated server (information processing device), database server, artificial intelligence (AI) system, number of communication terminals, sequence diagram, flowchart, images, notification content, and display screen are examples only, and the present invention is not limited to these.
[0018] Traditionally, building management (BM) has included building security, with security guards conducting patrols. These guards patrol the building, for example, every hour, checking for any abnormalities such as suspicious objects, suspicious individuals, lost items, signs of fire, or areas requiring cleaning. However, patrols by security guards can be inconsistent in their criteria for identifying abnormalities and their selection of areas to check. There is also a desire to reduce the workload of security guards and, consequently, lower building management costs. As described in Patent Document 1, some security systems using surveillance cameras acquire images captured by the cameras and generate monitoring results based on those images using pre-trained monitoring logic. However, existing technologies like those in Patent Document 1, while capable of detecting abnormal human behavior through machine learning, struggle to detect stationary objects, particularly those below a certain size.
[0019] In contrast, according to one embodiment of the present disclosure, events occurring at the security target may be detected by using generative AI, which has seen remarkable development in recent years, to find differences between time-series images from a surveillance camera. Furthermore, a notification regarding the latest status of the security target, including the events that have occurred, may be generated. This reduces the burden on security guards on their security duties and allows for the rapid detection of anomalies. Moreover, according to one embodiment of the present invention, the content of the changes that have occurred at the security target may be output in text form. This allows security guards to accurately grasp events that are difficult to judge from surveillance camera images alone.
[0020] <System Configuration> FIG. 1 is a diagram showing a configuration example of a security support system according to an embodiment of the present disclosure. The security support system 600 may be a system that supports security operations using a generative AI. Hereinafter, a case where the security support system 600 is applied to a building will be described, but the object to which the present disclosure is applied is not limited thereto. The security support system 600 may be applied to any building or area where security operations and management operations by security guards and managers (hereinafter also simply referred to as "security guards") are performed. The security support system 600 may be applicable to, for example, stores, accommodation facilities, condominiums, schools, entertainment facilities, event venues, etc.
[0021] The security support system 600 may include an imaging device (monitoring camera) 10, an information processing device (server) 100, a communication terminal (user terminal) 200 of a user such as a security guard or a manager, a generative AI system 300, and a database server 400. These may be connected to each other via a network 500 and may be capable of transmitting and receiving data.
[0022] The monitoring camera 10 is installed at a position where it can image the security target inside the building and may transmit an image to the server 100. The security target may be any place where security guards make their rounds. The security target may be, for example, an office, a rest room, an emergency staircase, an entrance / exit of the building, a corridor, etc. Although only one monitoring camera 10 is shown in FIG. 1, there may be a plurality of monitoring cameras 10 according to the security target. The monitoring camera 10 may also be referred to as a security camera.
[0023] Server 100 may perform various processes related to the security support system 600. These various processes related to the security support system may include, for example, the process of notifying the user terminal 200 in text about the status of the security target based on images of the security target acquired from the surveillance camera 10. For example, Server 100 may display a notification on the user terminal 200 regarding the latest status of the security target, based on the result of comparing the latest image from the surveillance camera 10 with past images. Here, Server 100 may use a generation AI to generate the above notification. Details of the processes of Server 100 will be described later.
[0024] The user terminal 200 may be a communication terminal for a security guard performing security duties. The user terminal 200 may receive notifications from the server 100 regarding the latest status of the security target and display them on the screen. The security guard may determine from the content of the notification which areas require particular patrolling. In Figure 1, a laptop computer is shown as the user terminal 200, but the user terminal 200 may be any terminal that can implement the functions described in each embodiment.
[0025] The generative AI system 300 may be a system that provides the functionality of a generative AI model via the network 500. In one embodiment of the present invention, the generative AI model provided by the generative AI system 300 may be a multimodal generative AI model. A multimodal generative AI model is an artificial intelligence system constructed by integrally processing multiple different types of information, such as text, images, and audio, using deep learning. Examples of multimodal generative AI models include "Gemini®" by Google®, "GPT-4®" or "GPT-4V®" by OpenAI, or "DALL E®". Alternatively, the generative AI system 300 may be "Azure OpenAI Service," which makes AI models developed by OpenAI available on Microsoft Azure® by Microsoft®. The execution platform for the generative AI model may be provided by the server 100.
[0026] The generation AI system 300 may perform processing based on prompts sent from the server 100. A prompt may be text for inputting instructions or questions to the AI. The prompts sent from the server 100 may be written in a predetermined programming language depending on the form of the generation AI system 300.
[0027] The database server 400 may store various types of information (data) used by the security support system 600. In Figure 1, only one database server 400 is shown separately from the server 100, but it may be integrated with the server 100. The database server 400 may, for example, store prompts input to the generation AI system 300. The database server 400 may also temporarily store images from the surveillance camera 10 and outputs from the generation AI system 300.
[0028] Network 500 may include wireless networks and wired networks. Specifically, for example, network 500 may be a wireless LAN (WLAN), a wide area network (WAN), LTE (Long Term Evolution), 4th generation communication (4G), 5th generation communication (5G), and 6th generation communication (6G) or later mobile communication systems. However, network 500 is not limited to these examples and may also include, for example, Bluetooth (registered trademark) or optical fiber lines. Network 500 may also be a combination of these.
[0029] <Security support processing> A security support process according to one embodiment of the present invention will be explained using Figures 2 and 3. Figure 2 may be an example of a sequence between a server 100, a surveillance camera 10, a generation AI system 300, and a user terminal 200. Figure 3(a) may be an example of an image from the surveillance camera 10, and Figure 3(b) may be an example of a notification displayed on the user terminal 200.
[0030] Prior to step S10 in the sequence shown in Figure 2, the server 100 may acquire images of the security target under normal conditions, captured by the surveillance camera 10, as reference images. Normal conditions refer to a state in which no abnormalities have occurred at the security target. A state in which no abnormalities have occurred may refer to a state in which no action by security guards is required, such as a state in which no lost items or suspicious objects are left unattended, equipment is in its normal position, there is no dirt or trash, adequate brightness is maintained, and there are no persons in need of rescue.
[0031] Figure 3(a) shows an example of a reference image. Figure 3(a) may be an image taken by a surveillance camera 10 installed in a break room, as an example. The reference image 30 may be captured in advance by the surveillance camera 10 and stored in the database server 400.
[0032] Server 100 may request an image from the surveillance camera 10 (step S10 in Figure 2). The surveillance camera 10 may send an image to Server 100 upon receiving the request from Server 100 (step S11). Server 100 may request the transmission of the image using a predetermined protocol according to the specifications of the surveillance camera 10. Server 100 may store the image acquired from the surveillance camera 10 (step S12). The image may be stored in the database server 400.
[0033] Server 100 may determine whether a predetermined time has elapsed since the last image was acquired (step S13). The predetermined time may be the interval at which security guards conventionally conducted patrols, such as one hour or two hours, but is not limited to these. To determine whether the predetermined time has elapsed, a timer may be used, or the time at which the last image was acquired may be recorded, and the elapsed time from that time to the current time may be calculated internally to determine whether the predetermined time has elapsed.
[0034] If it is determined that a predetermined time has not elapsed since the last image was acquired (NO in step S13), the server 100 may wait for the predetermined time to elapse. If it is determined that a predetermined time has elapsed since the last image was acquired (YES in step S13), the server 100 may request an image from the surveillance camera 10 (step S14). The surveillance camera 10 may receive the request from the server 100 and send the image to the server 100 (step S15).
[0035] Server 100 may store the previously acquired image as a past image and the currently acquired image as the latest image (step S16). That is, a past image may be an image that is positioned immediately before the latest image in the time series. Here, images 31 and 32 in Figure 3(a) may be examples of a past image and the latest image, respectively. Server 100 may send the reference image 30, the past image 31, the latest image 32, and a prompt to the generating AI system 300 (step S17). Here, the prompt may include an instruction to output the result of comparing the reference image, the latest image, and the past image. Note that sending a prompt to the generating AI system 300 may mean that the prompt is input to the generating AI model. Note that if the data of the previously acquired image is corrupted or has been deleted due to some error, Server 100 may use an image acquired two steps prior to that as a past image. That is, Server 100 may use an image acquired before the latest image as a past image.
[0036] The generating AI system 300 may compare multiple images and determine the type of difference according to the prompt (step S18). In one embodiment of the present invention, the prompt may be based on the idea that an anomaly at the present time is detected by comparing the latest image 32 with the reference image 30, and the duration of the anomaly is determined by comparing the past image 31 with the latest image 32. An anomaly may refer to a state that is different from normal. This will be explained using Figure 3(a).
[0037] First, in the generation AI system 300, the generation AI model may compare at least two images from the reference image 30, past image 31, and latest image 32, and detect differences. Then, the generation AI system 300 may determine the type of difference between each image according to the input prompt. For example, the reference image 30 and past image 31 differ in that past image 31 contains a plastic bottle 21. Also, the reference image 30 and latest image 32 differ in that latest image 32 contains a plastic bottle 21, a file 22, and a person 23. Furthermore, past image 31 and latest image 32 differ in that latest image 32 contains a file 22 and a person 23. Based on these comparison results, the type of difference, the plastic bottle 21, may be determined by the prompt to be "a lost item that has existed for a long time." Also, the type of difference, the file 22, may be determined by the prompt to be "a lost item that has existed for a short period of time." Furthermore, the prompt may include characteristics of a person to determine if they are a suspicious person, and if the person 23, which is a difference, is determined to be a "person," it may be determined whether or not that person corresponds to a "suspicious person." In addition, the prompt may include instructions to determine if there is a difference when the position of an object included in the reference image is different in the latest image, and to determine the period during which the object's position is different based on the comparison result between the past image and the latest image. The types of differences included in the prompt may include, but are not limited to, the above-mentioned "lost item that has been there for a long time," "lost item that has been there for a short time," "suspicious person," "lost equipment," "moved equipment," "person in need of rescue," "burnt-out light bulb," "fire," "dirt," etc.
[0038] Returning to Figure 2, the generating AI system 300 may output the result of executing the prompt (step S19). The output from the generating AI system 300 may include the content of the differences and the type of the differences as a result of comparing each image. In the example in Figure 3(a), the content of the differences may be output as "lost item that has been there for a long time," "lost item that has been there for a short time," and "suspicious person," respectively. The server 100 may output a notification to the user terminal 200 based on the output of the generating AI system 300 (step S20). That is, the server 100 may output a notification to the user terminal 200 in which the content of the differences between the images and the type of the differences are described in text. The user terminal 200 may display the notification on its screen (step S21). That is, the server 100 may notify the user terminal 200 of the content of the text describing the differences in the latest state of the security target from the normal state.
[0039] Figure 3(b) shows an example of a notification displayed on the user terminal 200. Notification 40 may include information about suspicious persons, lost items that have been there for a long time, and lost items that have been there for a short time, along with text describing their contents. Notification 40 may also include the latest image 32. This allows the user to accurately understand the latest status. Notifications may be sent by email or by a designated communication tool. Notifications may also be displayed on the web, and a URL (Uniform Resource Locator) or website address may be sent to the user terminal 200.
[0040] Returning to Figure 2, the server 100 may delete past images from the database server 400 (step S22). Then, returning to step S13, after a predetermined time has elapsed, images may be acquired from the surveillance camera 10 (steps S14, S15). In the following step S16, the latest image 32 in Figure 3(a) becomes a past image, and the image acquired in the most recent step S15 becomes the latest image, and the above process may be repeated.
[0041] Thus, according to one embodiment of the present invention, the state of the security target can be determined from the differences between images from a surveillance camera 10 that are consecutive in time series, and the state of the security target may be notified in text to the communication terminal of the security guard monitoring the security target. Since the differences between images are extracted by a generative AI model, it becomes possible to detect small, stationary objects, which machine learning models are not good at. In addition, since the differences between images, i.e., events that occurred to the security target, are output as text, security guards can accurately grasp events that are difficult to judge from surveillance camera images alone.
[0042] Furthermore, when new images of the subject of security are acquired, past images may be deleted from the database server 400. This ensures the protection of personal information.
[0043] The server 100 may also obtain information regarding the location of the surveillance camera 10. This information can be obtained, for example, from the MAC address of the surveillance camera 10. The server 100 may include this location information in the notification sent to the user terminal 200. For example, as shown in Figure 3(b), the notification 40 may indicate that it pertains to the "10th floor break room camera".
[0044] The prompt may include instructions in the notification to list the types of differences based on the priority order set according to the type of difference. Priority may be set higher for events that require immediate attention by security guards. For example, priorities are not limited to these, but may be set as follows: "person in need of rescue," "fire," "suspicious person," "lost item that has been there for a long time," "lost item that has been there for a short time," etc. This will enable security guards to quickly identify events that require immediate attention. In the example of notification 40 in Figure 3(b), the detection of a suspicious person is listed as the first item.
[0045] Furthermore, the office layout may be changed. In that case, images taken by the surveillance camera 10 during normal operation after the layout change may be stored as reference images. This will ensure the accuracy of security.
[0046] In the sequence shown in Figure 2, the latest image is described in a manner in which the latest image is acquired at predetermined intervals. However, the present invention is not limited to this, and the server 100 may acquire the latest image when predetermined conditions are met. These predetermined conditions may be, for example, when a request is made by the user, or when an image is transmitted by the control of the surveillance camera 10. Furthermore, the frequency of acquiring the latest image may differ depending on the installation location of the surveillance camera 10. This makes it possible to provide a security support system that can respond to user requests more flexibly.
[0047] <Structure> Figure 4 will be used to explain the hardware and functional configuration of server 100.
[0048] <server> (1) Server hardware configuration The server 100 may include a control unit 110, a communication unit 120, an input / output unit 130, and a storage unit 170.
[0049] The control unit 110 is typically a processor, and may include a central processing unit (CPU), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), a microprocessor, etc., and may be implemented by logic circuits (hardware) or dedicated circuits formed on an integrated circuit chip (IC (Integrated Circuit) chip, LSI (Large Scale Integration)), etc.
[0050] The communication unit 120 may be implemented as hardware such as a NIC (Network Interface Card), a network adapter, communication software, or a combination thereof. The communication unit 120 may send and receive various types of data to and from the surveillance camera 10 and the user terminal 200 via the network 500.
[0051] The input / output unit 130 may include an input device for inputting various operations to the server 100, and an output device for outputting processing results processed by the server 100. The input device may include, for example, hardware keys such as a touch panel, touch display, or keyboard, a pointing device such as a mouse, a camera, or a microphone. The output device may output processing results processed by the control unit 110. The output device may include, for example, a display, touch panel, or speaker.
[0052] The storage unit 170 stores various programs and data necessary for the operation of the server 100. The storage unit 170 may include, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), flash memory, etc. The storage unit 170 may also include memory that provides a working area for the control unit 110.
[0053] (2) Server Functional Configuration The server 100 may include an image acquisition unit 111, an instruction unit 112, a result acquisition unit 113, a notification unit 114, and a reception unit 115, as functions implemented by the control unit 110. Note that some of the functional units shown in Figure 4 may be omitted if they are not essential in each embodiment. Furthermore, the functions or processing of each functional unit may be implemented by machine learning or AI to the extent feasible.
[0054] The image acquisition unit 111 may acquire images of the security target captured by the surveillance camera (imaging device) 10. The image acquisition unit 111 may also acquire images of the security target from the surveillance camera 10 when predetermined conditions are met. These "predetermined conditions" may include the elapsed time since the last image acquisition. In other words, the image acquisition unit 111 may acquire images at predetermined intervals.
[0055] Furthermore, the image acquisition unit 111 may store a reference image, which is a reference image set in advance as a reference image from among the images of the security target, the latest image, which is the most recent image in the time series, and past images, which are images older than the latest image, in a predetermined storage unit. The predetermined storage unit may be a database server 400. Here, the image acquisition unit 111 may delete past images from the storage unit that are not used for comparison with the newly acquired latest image when an image of the security target is newly acquired. Furthermore, the image acquisition unit 111 may acquire information regarding the installation location of the surveillance camera 10.
[0056] The instruction unit 112 may input the reference image, the latest image, and past images to the generation AI model. Inputting data to the generation AI model may refer to sending data to the generation AI system 300. The instruction unit 112 may further input a prompt to the generation AI that includes an instruction to output the result of comparing the reference image, the latest image, and past images. Here, the instruction unit 112 may use the image located immediately before the latest image in the time series as the past image. Furthermore, the instruction unit 112 may input an instruction to the generation AI model that outputs the type of difference between the reference image, the latest image, and past images based on the result of comparing them. In addition, if multiple differences are detected, the instruction unit 112 may input a prompt to the generation AI model that includes an instruction to list the types of differences in order of priority set according to the type of difference.
[0057] The instruction unit 112 may also issue an instruction to output the result of comparing the reference image, the latest image, and the past image, and may input the reference image, the latest image, and the past image to the generating AI model at a different timing than the prompt. In this way, by prioritizing the analysis and understanding of images by the generating AI model, it is possible to respond quickly when a prompt is input. In addition, by inputting the images in advance, it is not necessary to input images each time multiple different prompts are input.
[0058] The result acquisition unit 113 may acquire the execution results of prompts output by the generating AI model. The result acquisition unit 113 may acquire the output of the generating AI model in a predetermined data format. The predetermined data format may be a predetermined communication tool or communication application used on the user terminal 200, and may be, but is not limited to, JSON, XML (Extensible Markup Language), HTML (HyperText Markup Language), etc.
[0059] The notification unit 114 may generate a notification based on the output of the generated AI model. The notification unit 114 may also output the notification to the communication terminal of the user monitoring the security target. The notification unit 114 may also generate the notification including the latest image.
[0060] <Server control flowchart> The control method for the server 100 described above will be explained using the flowchart in Figure 5. First, the image acquisition unit 111 of the server 100 may acquire images of the security target captured by the imaging device (surveillance camera) (step T11). Next, the instruction unit 112 may input a prompt to the generating AI model that includes an instruction to output in text the result of comparing the images of the security target with a reference image that has been set in advance as a reference image, the latest image which is the most recent image in the time series, and past images which are images older than the latest image (step T12). The notification unit 114 may generate a notification based on the output of the generating AI model (step T13).
[0061] The present invention has been described based on various drawings and embodiments, but it should be noted that those skilled in the art will find it easy to make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of the present invention. For example, the functions included in each component, step, etc., can be rearranged in a logically consistent manner, and multiple components or steps, etc., can be combined into one or divided. Furthermore, the configurations shown in the above embodiments may be combined as appropriate. For example, each component described as being provided by server 100 may be implemented by multiple servers in a distributed manner. Furthermore, the functions described as being performed by server 100 may be performed by the generation AI system 300. Also, the functions described as being performed by the generation AI system 300 may be performed by server 100.
[0062] For example, the above description explains a mechanism in which a prompt is input to the generating AI model in response to the acquisition of the latest image. However, if no anomaly is detected, the prompt does not need to be executed. For example, if the server 100 is equipped with an image analysis unit and no difference is detected between the latest image and the reference image through image analysis processing, the prompt does not need to be input to the generating AI model. This reduces the resources required to use the generating AI system.
[0063] Furthermore, although only one server 100 is shown in Figure 1, it is not limited to this. That is, each function described as being provided by server 100 may be implemented by multiple servers. Also, server 100 may be a distributed server system that operates cooperatively by communicating over a network, for example, or it may be a so-called cloud server. In other words, server 100 may include not only physical servers but also virtual servers created by software. Moreover, each function described as being performed by server 100 may be provided by a cloud-based platform that operates over network 500.
[0064] Furthermore, although the user terminal 200 is shown as a laptop computer in Figure 1, the user terminal 200 is not limited to this and may be a computer (e.g., a smartphone, tablet, or desktop computer) or a wearable device (such as glasses or a watch).
[0065] Furthermore, although only one database server 400 is shown separately from server 100 in Figure 1, it may be integrated with server 100. That is, the database server 400 may be the volatile or non-volatile memory of server 100. Also, the database server 400 may consist of multiple storage devices. In addition, the database server 400 may be connected to server 100 via a dedicated internal network different from network 500.
[0066] Furthermore, as described above, the prompts may be set in multiple stages. For example, a first prompt may be provided to compare a reference image, a past image, and the latest image, and a second prompt may be provided to determine the type of difference between the images if there is a difference based on the comparison result of the first prompt. In other words, there may be one prompt or two or more prompts sent to the generating AI model.
[0067] The programs of each embodiment of this disclosure may be provided stored in a storage medium readable by the information processing device. The storage medium is a “non-temporary tangible medium” capable of storing programs. The programs include, for example, software programs and information processing device programs. When each functional unit of the server 100 as an information processing device is implemented by software, the server 100 functions as an image acquisition unit 111, an instruction unit 112, a result acquisition unit 113, a notification unit 114, and a reception unit 115 by having the processor execute a program loaded into memory.
[0068] The storage medium may, where appropriate, include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable storage medium, or two or more suitable combinations thereof. The storage medium may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.
[0069] Furthermore, each embodiment of this disclosure can also be realized in the form of a data signal embedded in a carrier wave, where the program is embodied by electronic transmission. The programs of this disclosure may be implemented using, for example, scripting languages such as JavaScript® and Python®, or languages such as C, Go, Swift®, Koltin®, and Java®.
[0070] According to each aspect of this disclosure described above, by providing users with a more user-friendly work environment, we can contribute to achieving Sustainable Development Goal (SDG) 11, "Make cities and human settlements inclusive, safe, resilient and inclusive, safe [Explanation of symbols]
[0071] 10. Surveillance camera (imaging device) 100 Servers (Information Processing Devices) 110 Control Unit 111 Image acquisition unit 112 Instruction section 113 Result acquisition part 114 Notification Department 115 Reception Department 120 Communications Department 130 Input / output section 170 Storage section 200 User terminals (communication terminals) 300 Generating AI Systems 400 Database Servers 500 Networks 600 Security Support System
Claims
1. An image acquisition unit that acquires images of the security target captured by the imaging device, A prompt that inputs a prompt to the generating AI model, which includes instructions to output in text the results of comparing a combination of a reference image (pre-set as a reference image) from the images of the subject to security, the latest image (the most recent image in the time series), and past images (images acquired before the latest image), A notification unit that generates notifications based on the output of the generated AI model, Equipped with, The instruction unit outputs the type of difference between the reference image, the latest image, and the past image based on the results of comparing the reference image, the latest image, and the past image, and if multiple differences are detected, it inputs the prompt to the generating AI model, which includes an instruction to enumerate the types of differences in order of priority set according to the type of difference.
2. The instruction unit, in the instruction, designates the image located immediately before the latest image in the time series as the past image. The information processing apparatus according to claim 1.
3. The image acquisition unit acquires an image of the security target from the imaging device when a predetermined condition is met. The instruction unit inputs the prompt to the generating AI model in response to the acquisition of a new image of the security target. The information processing apparatus according to claim 1 or 2.
4. The image acquisition unit acquires an image of the security target from the imaging device when a predetermined time has elapsed since the previous image was acquired, as a predetermined condition. The information processing apparatus according to claim 3.
5. The notification unit outputs the notification to the communication terminal of the user monitoring the security target. The information processing apparatus according to claim 1.
6. The system further includes a reception unit that accepts the selection of an image from among the images of the subject to security to be set as the reference image. The information processing apparatus according to claim 1.
7. The instruction unit provides instructions to output the result of comparing the reference image, the latest image, and the past image, and inputs the reference image, the latest image, and the past image to the generating AI model at different timings. The information processing apparatus according to claim 1.
8. The notification unit generates the notification including the latest image. The information processing apparatus according to claim 1.
9. Information processing device, The steps include acquiring an image of the security target captured by the imaging device, The steps include inputting a prompt to a generating AI model that includes instructions to output a textual result of comparing a combination of a reference image (pre-set as a reference image from among the images of the subject to security), the latest image (the most recent image in the time series), and past images (images acquired earlier in the time series than the latest image), and The steps include generating a notification based on the output of the generated AI model, Execute, The input step is a control method for an information processing device, which involves inputting a prompt to the generating AI model, which includes an instruction to output the type of difference between the reference image, the latest image, and the past image based on the results of comparing the reference image, the latest image, and the past image, and if multiple differences are detected, to enumerate the types of differences in order of priority set according to the type of difference.
10. In an information processing device, A function to acquire images of the security target captured by the imaging device, A function to input prompts into a generating AI model that include instructions to output in text the results of a comparison between a pre-set reference image, the most recent image in the time series, and past images acquired earlier in the time series than the latest image, from among the images of the subject to security. A function to generate notifications based on the output of the aforementioned generated AI model, To make it happen, The input function is a control program for an information processing device that, based on the results of comparing the reference image, the latest image, and the past image, outputs the type of difference between the reference image, the latest image, and the past image, and if multiple differences are detected, inputs a prompt to the generating AI model that includes an instruction to enumerate the types of differences in order of priority set according to the type of difference.
Citation Information
Patent Citations
Still image recorder
JP1999261855A
Terminal unit
JP2005328189A
Automatic suspicious object detecting device
JP2009187348A
Object recognition system, monitoring system using the same, and watching system
JP2011198244A
Information processing apparatus, information processing method, and camera
JP2021012657A