Image display system and image display program

A dual-stage image selection process on edge and cloud devices addresses resource constraints by performing initial rough selection on edge devices and detailed selection on clouds, optimizing communication and processing loads while ensuring accurate image display.

JP7732695B1Active Publication Date: 2025-09-02AWL INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025001868
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-09-02
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Existing image display systems face challenges in reducing communication and processing loads on edge devices due to the need to transfer and analyze all camera images, which is resource-intensive and inefficient.

Method used

Implementing a dual-stage selection process where edge devices perform primary selection based on rough criteria and cloud servers perform secondary selection based on detailed criteria, reducing the amount of data transferred and processing load on edge devices.

Benefits of technology

This approach effectively reduces communication and processing loads on edge devices while ensuring accurate selection and display of user-desired images, enhancing system efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007732695000001_ABST
    Figure 0007732695000001_ABST
Patent Text Reader

Abstract

To provide an image display system and an image display program capable of reducing the amount of communication between an edge device side and a cloud side and also reducing the processing load on the edge device. [Solution] In an image display system that performs each function, an edge device performs a primary selection by detecting whether an object is captured in a frame image captured by a camera according to rough selection criteria. The edge device selects frame images in which an object is detected and transmits them to a cloud server, and the cloud server performs a secondary selection by selecting frame images in which the object is captured according to detailed selection criteria from the selected frame images. The selected frame images are stored in a storage unit so that they can be viewed from a client 4. This narrows down the frame images to those most likely to capture the object desired by the user and transmits them to the cloud server.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image display system and an image display program. [Background technology]

[0002] Conventionally, by using a VMS (Video Management System), images captured by cameras such as surveillance cameras installed in facilities such as stores can be transferred to and stored on a server on the cloud, allowing users such as operators in the management department to view them on display devices in remote locations.

[0003] There is also a known system that performs image analysis (object detection and object recognition) on images captured by a camera installed in a facility such as a store using a device (a so-called edge device) located on the facility side where the camera is installed (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-88157 Summary of the Invention [Problem to be solved by the invention]

[0005] As described above, a VMS allows users to view images captured by cameras such as surveillance cameras installed in a facility from a remote display device. However, users may not want to view all of the images captured by the camera, but only a portion of the images (for example, a video of a customer engaging in suspicious behavior). However, if all of the images captured by the camera were transferred to and stored on a cloud server, the amount of communication between the edge device and the server (cloud side) would increase. Therefore, it is conceivable to use an image analysis system on the edge device, such as that described in Patent Document 1, to detect the target that the user wants to view, and select only frame images that show the target that the user wants to view and transfer them to a cloud server.

[0006] However, due to the limited computer resources of edge devices, it is difficult for the edge device alone to perform the above-described process of detecting the user's desired viewing subject (the process of selecting a frame image showing the user's desired viewing subject).

[0007] The present invention aims to solve the above-mentioned problems and to provide an image display system and an image display program that can reduce the amount of communication between the edge device side and the cloud side and also reduce the processing load on the edge device. [Means for solving the problem]

[0008] In order to solve the above-mentioned problems, an image display system according to a first aspect of the present invention includes a primary selector unit that performs a primary selection of frame images on the edge device side, which is a process of detecting a viewing target desired by a user from frame images input from a camera according to rough selection criteria, and selecting frame images in which the viewing target may appear; ,before an edge device-side transmitting unit that transmits image information corresponding to the frame image selected by the primary selector unit from the edge device side to the cloud side; A picturea cloud-side storage unit for storing image information; and A picture a secondary selector unit that performs secondary selection of frame images, which is a process of selecting frame images showing a user-desired viewing target in accordance with detailed selection criteria using image information; and a cloud-side selection unit that selects frame images showing a user-desired viewing target in accordance with detailed selection criteria using the image information. Selection The time information corresponding to the selected frame image is used to select a frame image in which the user's desired viewing subject is displayed. movie of, Reading from the storage device on the edge device side, and an image display control unit that controls the image to be displayed on the cloud-side display unit.

[0009] In this image display system, it is desirable that the primary selector section outputs, as a result of the primary selection, text information relating to a frame image in which the viewing subject may appear.

[0010] In this image display system, it is desirable that the edge device side transmitting unit transmits to the cloud side text information regarding a frame image in which the object to be viewed may appear and time information when the frame image in which the object to be viewed may appear was taken.

[0011] In this image display system, the image information transmitted by the edge device side transmission unit and stored in the cloud side storage unit is the text information output by the primary selector unit, and it is desirable that the secondary selector unit uses the text information to perform secondary selection of the frame image.

[0014] In this image display system, a selection operation unit is provided for a user to select a desired image from the images displayed on the cloud-side display unit, and an image selected by the user using the selection operation unit is provided. Videos containing from the storage device on the edge device side and store it in a storage device of a device other than the edge device.

[0015] In this image display system, the secondary selector section may perform secondary selection of the frame images using a VLM (Visual Language Model) or an LLM (Large Language Model).

[0016] An image display program according to a second aspect of the present invention includes a computer including a primary selector unit that performs a primary selection of frame images, which is a process of detecting a viewing target desired by a user from frame images input from a camera on the edge device side according to rough selection criteria, and selecting frame images that may include the viewing target; ,before an edge device-side transmitting unit that transmits image information corresponding to the frame image selected by the primary selector unit from the edge device side to the cloud side; A picture a cloud-side storage unit for storing image information; and A picture a secondary selector unit that performs secondary selection of frame images, which is a process of selecting frame images showing a user-desired viewing target in accordance with detailed selection criteria using image information; and a cloud-side selection unit that selects frame images showing a user-desired viewing target in accordance with detailed selection criteria using the image information. Selection The time information corresponding to the selected frame image is used to select a frame image in which the user's desired viewing subject is displayed. movie of, Reading from the storage device on the edge device side, It functions as an image display control unit that controls the display on the cloud side. [Effects of the Invention]

[0017] According to the image display system according to the first aspect of the present invention and the image display program according to the second aspect, unlike conventional image display systems using a VMS, only frame images selected by a primary selector or image information corresponding to the frame images selected by the primary selector are transmitted from the edge device to the cloud, rather than all images captured by the camera. This reduces the amount of communication between the edge device and the cloud. Furthermore, in the process of selecting frame images showing a user's desired viewing target, only the process of detecting the user's desired viewing target from frame images input from the camera according to rough selection criteria (primary selection of frame images) is performed on the edge device, while the process of selecting frame images showing the user's desired viewing target according to detailed selection criteria (secondary selection of frame images) is performed on the cloud. This reduces the processing load on the edge device and allows images showing the user's desired viewing target to be accurately selected and displayed on the cloud display. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a schematic diagram of an image display system according to a first embodiment. [Figure 2] FIG. 2 is a block diagram showing a configuration of an edge device. [Figure 3] FIG. 2 is a block diagram showing the configuration of a cloud server. [Figure 4] FIG. 2 is a block diagram showing the configuration of a client; [Figure 5] FIG. 1 is a functional block diagram of an image display system. [Figure 6] 10 is a flowchart illustrating an example of a processing procedure in the image display system. [Figure 7] 10 is a flowchart illustrating an example of a processing procedure in the image display system. [Figure 8] 10 is a flowchart illustrating an example of a display control procedure in the image display system. [Figure 9]FIG. 10 is an explanatory diagram showing an example of an image displayed on a client. [Figure 10] FIG. 10 is an explanatory diagram showing an example of an image displayed on a client. [Figure 11] FIG. 10 is a functional block diagram of an image display system according to a second embodiment. [Figure 12] 10 is a flowchart showing an example of a processing procedure in the image display system of the second embodiment. [Figure 13] 10 is a flowchart showing an example of a processing procedure in the image display system of the second embodiment. [Figure 14] 10 is a flowchart showing an example of a processing procedure performed by a selection operation unit and an image display control unit according to the second embodiment. [Figure 15] FIG. 10 is a functional block diagram of an image display system according to a third embodiment. [Figure 16] FIG. 10 is a functional block diagram of an image display system according to a fourth embodiment. [Figure 17] 10 is a flowchart showing an example of a processing procedure in an image display system according to a fourth embodiment. [Figure 18] 10 is a flowchart showing an example of a processing procedure in an image display system according to a fourth embodiment. [Figure 19] 13 is a flowchart showing an example of a processing procedure performed by a selection operation unit and an image display control unit according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0019] The present disclosure will be specifically described with reference to the drawings showing embodiments thereof.

[0020] [First embodiment] 1 is a schematic diagram of an image display system 100 according to a first embodiment. The image display system 100 includes a camera 2 installed in a target space such as a store, an edge device 1 connected to the camera 2, a cloud server 3 capable of communication with the edge device 1, and a client 4 capable of communication with the cloud server 3. A plurality of edge devices 1 and a plurality of cameras 2 may be installed in the same space.

[0021] The camera 2 uses an image element that is responsive to visible light and / or near-infrared light and outputs image information. The camera 2 outputs image information in time series at a rate of several fps to several tens of fps. The camera 2 is installed so as to look down from the top of the space, such as on the ceiling or shelf of the space in which it is installed. The camera 2 may be a type that is attached to the ceiling and has a field of view that covers the entire space 360 ​​degrees. The camera 2 sequentially transmits image information to the edge device 1 via the local network LN.

[0022] The camera 2 and the edge device 1 can be connected to each other via a local network LN, which may be wireless or wired. The local network LN may be a wired LAN or a wireless network such as WiFi or Bluetooth (registered trademark).

[0023] The edge device 1 can be connected to the cloud server 3 via a network N. The network N is a wired or wireless communication network that may include a public communication network, a dedicated line, or a carrier network. The client 4 can be connected to the cloud server 3 and the edge device 1 via the network N.

[0024] The cloud server 3 stores in a database 310 a group of learning models used in processes selectively executed by the edge device 1 and the cloud server 3. The database 310 may also include a group of learning models provided by an external service outside the system. The cloud server 3 reads out from the database 310 a learning model corresponding to an object to be detected by the processes executed by the edge device 1 and the cloud server 3, and deploys it to the edge device 1 and the cloud server 3.

[0025] The image display system 100 of the first embodiment extracts feature amounts from an image captured by a camera 2 installed in a target space, detects an object such as a person or object to be detected from the image based on the feature amounts, recognizes attributes of the detected object, and outputs the recognition result. The image display system 100 is a system that can automatically display on a client 4 an image showing an object desired by a user.

[0026] The image display system 100 of the first embodiment is a system that allows a user to select the attribute of a person or object to be detected from an image. Furthermore, the image display system 100 of the first embodiment reduces the amount of communication between the edge device 1 and the cloud server 3, and further reduces the processing load on the edge device 1 by sharing processing between the edge device 1 and the cloud server 3.

[0027] The configurations of the edge device 1, cloud server 3, and client 4 for realizing such an image display system 100, as well as the processing executed by each of them, will be described in detail below.

[0028] 2 is a block diagram showing the configuration of the edge device 1. The edge device 1 is a box-shaped device that can be installed in a target space together with a camera 2. The edge device 1 includes a processing unit 10, a storage unit 11, a first communication unit 12, and a second communication unit 13.

[0029] The processing unit 10 includes one or more processors such as a central processing unit (CPU), a micro-processing unit (MPU), a graphics processing unit (GPU), or a neutral processing unit (NPU). The processing unit 10 includes a memory that is a temporary storage medium such as a static random access memory (SRAM) or a dynamic random access memory (DRAM). The processing unit 10 includes a timer and can acquire time information at each point in time from data from the timer. The processing unit 10 may be configured as a single piece of hardware (SoC: System On a Chip) that integrates a processor, a memory, a storage unit 11, a first communication unit 12, and a second communication unit 13. The specifications of the processing unit 10 may be the same or different between edge devices 1.

[0030] The processing unit 10 reads the first image display program P1 stored in the storage unit 11 into the memory and executes it, thereby causing the processor to execute various processes described below and function as the edge device 1 of the present disclosure.

[0031] The storage unit 11 is a relatively large-capacity non-transitory storage medium such as a hard disk, a flash memory, etc. A part of the storage unit 11 may be removable.

[0032] The storage unit 11 stores a program (program product) required for the processing unit 10 to execute processing, the results of processing by the processing unit 10, and reference setting information. The setting information includes an identifier for the edge device 1, identification data for the connected camera 2, etc. The program product includes an OS (Operating System) program, a first image display program P1 that runs on the OS in the edge device 1, and a learning model group M1. The learning model group M1 will be described in detail later.

[0033] The first image display program P1 or the learning model group M1 stored in the memory unit 11 may be the first image display program P9 and the learning model group M9 stored in a computer-readable non-temporary storage medium 9 that are read by the processing unit 10 and stored in the memory unit 11, or may be pre-stored at the time of shipment. The first image display program P1 or the learning model group M1 stored in the memory unit 11 may be downloaded by the processing unit 10 from the database 310 of the cloud server 3 via the second communication unit 13 or from another download server.

[0034] The storage unit 11 stores image information acquired from the connected camera 2 in chronological order in association with time information. The image information may be frame images or moving images. The storage unit 11 stores only image information in which a user's desired object (viewing target) is detected as a result of processing by the processing unit 10, which will be described later. The storage unit 11 may continue to store data by overwriting it according to a FIFO (First In First Out) method within a range appropriate to the capacity. The storage unit 11 may also store text information describing the detected object, text information indicating the attributes of the detected object, etc. in association with time information.

[0035] The first communication unit 12 is a communication device that realizes communication via a local network LN installed in the target space. The first communication unit 12 may be a LAN network card or a CAN communication device. The first communication unit 12 may be a communication device compatible with wireless networks such as WiFi or Bluetooth (registered trademark). The first communication unit 12 may include multiple communication devices compatible with various types of cameras 2. The first communication unit 12 may include an interface such as a USB (Universal Serial Bus) connected to the camera 2. The first communication unit 12 can be replaced by an interface connected to the camera 2 via a coaxial cable or another serial bus. The processing unit 10 acquires image information from the camera 2 via the local network LN using the first communication unit 12. The first communication unit 12 may be the same device as the second communication unit 13 described below.

[0036] The second communication unit 13 is a communication device that realizes communication via an external network N. The second communication unit 13 may be a network card for a wired LAN, or may be a communication device that realizes carrier communication via a carrier network. The second communication unit 13 may be a communication device that supports a wireless network such as WiFI or Bluetooth (registered trademark). The second communication unit 13 may support secure communication with the cloud server 3 such as SSL. The second communication unit 13 may be an interface that realizes a communication connection with the cloud server 3 via a dedicated line.

[0037] 3 is a block diagram showing the configuration of the cloud server 3. The cloud server 3 may be configured as a single server computer, or may be configured to distribute processing among multiple server computers. The cloud server 3 includes a processing unit 30, a storage unit 31, and a communication unit 32.

[0038] The processing unit 30 includes one or more processors such as a CPU, an MPU, a GPU, an NPU, etc. The processing unit 30 includes a memory that is a temporary storage medium such as an SRAM or a DRAM.

[0039] The storage unit 31 is a relatively large-capacity non-temporary storage medium such as a hard disk, flash memory, etc. The storage unit 31 stores programs (program products) and setting information required for the processing unit 30 to execute the processes described below.

[0040] The program products stored in the storage unit 31 include a server program P31. The server program P31 includes a module that functions as a data server that reads out a group of models stored in the database 310 and transmits them to the edge device 1. The server program P31 includes a module that functions as a web server, and can output the results of processing executed by the cloud server 3 to the client 4 via a web page.

[0041] The program products stored in the storage unit 31 include a second image display program P32 that runs on the cloud server 3, and a learning model group M3. The second image display program P32 is a program that causes a processor to execute processing linked to processing based on the first image display program P1 in the edge device 1. The learning model group M3 is selected from and stored in the database 310. The second image display program P32 and the learning model group M3 will be described in detail below.

[0042] The server program P31, the second image display program P32, or the learning model group M3 stored in the memory unit 31 may be the server program P81 and the second image display program P82 stored in a computer-readable non-transitory storage medium 8 that the processing unit 30 reads and stores in the memory unit 31. The learning model group M3 may also be the learning model group M8 that was stored in the non-transitory storage medium 8 that the processing unit 30 reads and stores in the memory unit 31. The server program P31 and the second image display program P32 stored in the memory unit 31 may be downloaded by the processing unit 30 from another download server via the second communication unit 13.

[0043] The setting information stored in the storage unit 31 includes data for identifying the edge device 1, which is associated with data for identifying the space in which the edge device 1 is installed, and a correspondence relationship between the data and the name. The cloud server 3 executes processes simultaneously and in parallel for multiple target spaces, and may store identification data of edge devices 1 or spaces for which user requests are permitted as a whitelist, associated with the user's account data, in the storage unit 31. This allows the cloud server 3 to identify the target edge device 1 when the user specifies the name of a space and specifies which object is to be detected in an image captured by the camera 2 installed in that space.

[0044] The storage unit 31 stores image information captured by the camera 2 and selected by primary selection or secondary selection. The image information may be frame images, chronologically consecutive frame images, or moving images separated into files in predetermined time units.

[0045] The database 310 may be built in the storage unit 31 or in an external storage device. As described above, part of the database 310 may include a model providing service used on the Web, which is connected to the database 310 via the network N. The database 310 holds a group of learning models used in processes executed by the edge device 1 and the cloud server 3. The group of learning models includes, for example, detection learning models such as a person detector, a head detector, a face detector, an animal detector, and a specific device detector. The group of learning models also includes recognition learning models that recognize attributes, such as an age recognizer, an eyeglasses wear recognizer, a hat wear recognizer, a face orientation recognizer, or a posture recognizer. The group of learning models includes large and small language models (LLMs, SLMs), vision and language models (VLMs), etc.

[0046] The communication unit 32 is a communication device that realizes a communication connection with the client 4 and the edge device 1 via the network N.

[0047] 4 is a block diagram showing the configuration of the client 4. The client 4 is a personal computer, a smartphone, or a tablet terminal. The client 4 may be used by the administrator of the space where the camera 2 is installed, or by the operator of the management company of the cloud server 3.

[0048] The client 4 includes a processing unit 40, a storage unit 41, a communication unit 42, a display unit 43, and an operation unit 44. The processing unit 40 includes one or more processors such as a CPU, an MPU, a GPU, or an NPU. The processing unit 40 includes a memory that is a temporary storage medium such as an SRAM or a DRAM.

[0049] The storage unit 41 is a memory of a non-temporary storage medium such as a hard disk or flash memory. The storage unit 41 stores the functions of the data server provided by the cloud server 3 and a client program P4 for the web server. The client program P4 is, for example, a web browser program. The client program P4 is a program that causes the processing unit 40 to execute a process of displaying various data, including images, provided by the cloud server 3 on a screen.

[0050] The communication unit 42 is a communication device that establishes a communication connection with the cloud server 3 via the network N. The communication unit 42 may be a communication device that establishes a communication connection with the cloud server 3 via a dedicated line. The communication unit 42 may be a communication device that establishes a direct communication connection with the second communication unit 13 of the edge device 1 via a wireless communication medium, a USB cable, or the like.

[0051] The display unit 43 uses a display such as a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 43 displays a web page including text and images through processing based on the client program P4 of the processing unit 40. The display unit 43 may use a display with a built-in touch panel.

[0052] The operation unit 44 is a user interface such as a keyboard or a pointing device that accepts operations from a user or an operator. The operation unit 44 may be a touch panel built into the display of the display unit 43, or may be physical buttons. The operation unit 44 may be a voice input unit that accepts operations by voice using a voice recognition function. The operation unit 44 can notify the processing unit 40 of operation information by the user or operator.

[0053] In the image display system 100 configured in this manner, the edge device 1 and the cloud server 3 share the processing of images taken by the camera 2 in each space, and the image information showing the object is displayed on the client 4 so that the user can view it.

[0054] Fig. 5 is a functional block diagram of the image display system 100. In the image display system 100, the functions shown in Fig. 5 are shared and performed based on a first image display program P1 in the edge device 1 and a second image display program P32 in the cloud server 3. In the image display system 100, the processing unit 10 in the edge device 1 functions as a primary selector unit 101 and an edge device-side transmission unit 102, and the processing unit 30 in the cloud server 3 functions as a cloud-side storage unit 301, a secondary selector unit 302, a selection operation unit 303, and an image display control unit 304.

[0055] In the first embodiment, the primary selector unit 101 receives frame images captured by the camera 2 via the first communication unit 12 and selects frame images that may contain the object desired by the user using a detection learning model M11 for the object. The detection learning model M11 is selected from human detectors, head detectors, face detectors, animal detectors, and specific device detectors provided in the database 310 according to the object to be detected set for each space, and is stored in the storage unit 11. Furthermore, an object detection process using the detection learning model M11 is deployed in the edge device 1. The primary selector unit 101 inputs frame images captured by the camera 2 to the detection learning model M11 and obtains detection results (coordinate data and accuracy of the area in which the object is captured) output from the detection learning model M11. If the detection results obtained from the detection learning model M11 indicate that an object is captured, the primary selector unit 101 selects the target frame image. The primary selector unit 101 stores the selected frame images in the storage unit 11 and may discard frame images not selected. The primary selector unit 101 stores a predetermined time period of video before and after the selected frame image in the storage unit 11 as image information. When the primary selector unit 101 selects chronologically consecutive frame images, it may store them in the storage unit 11 as a continuous video. The processing unit 10 may improve the detection accuracy of the detection learning model M11 by referring to the image size, etc. included in the configuration data. When the processing unit 10, as the primary selector unit 101, cannot detect an object such as a person, it ends processing for the target frame image and does not perform any further processing, but instead performs processing for the next frame image.

[0056] As a result of the primary selection, the primary selector unit 101 may store text information associated with the frame image regarding a frame image that may show an object. For example, if the detection target of the detection learning model M11 is a "person," a "head," a "face," an "animal," or a "specific device," the primary selector unit 101 may store "person," "head," "face," "animal," or "specific device" as text information describing the detected object.

[0057] In the first embodiment, the edge device-side transmitting unit 102 transmits the frame image selected by the function of the primary selector unit 101 to the cloud server 3 via the second communication unit 13. The edge device-side transmitting unit 102 may transmit a video including frames before and after the frame image to the cloud server 3, or may transmit text information (image information) including time information corresponding to the frame image to the cloud server 3. The time information is information indicating the time corresponding to the frame image, a timestamp, or the elapsed time or count from a specific time. If the text information has been stored by the primary selector unit 101 in association with the frame image, the edge device-side transmitting unit 102 may also transmit this text information to the cloud server 3.

[0058] The cloud-side storage unit 301 stores in the memory unit 31 the frame image or image information transmitted from the edge device 1 by the edge device-side transmission unit 102. The frame image or image information stored in the memory unit 31 is the frame image or image information selected by the edge device 1. When the cloud-side storage unit 301 receives not only the frame image but also time information or text information from the edge device 1, it also stores this time information or text information (image information) in the memory unit 31 in association with the identification data of the edge device 1, the identification data of the camera 2, etc.

[0059] The secondary selector unit 302 further selects frame images showing objects from the frame images or image information stored in the memory unit 31 by the cloud-side storage unit 301, in accordance with detailed selection criteria. The detailed selection criteria include criteria for attributes for determining whether an object in a frame image in which an object has been detected by the primary selector unit 101 is an object desired by the user. For example, if the primary selector unit 101 detects that a person is shown in the frame image, the detailed selection criteria are criteria for the attributes of the person. For example, if the primary selector unit 101 detects a person, the secondary selector unit 302 uses criteria such as whether the frame image shows a person in a specific posture or a person of a specific age.

[0060] Other examples are possible as long as the selection criteria in the primary selector unit 101 and the selection criteria in the secondary selector unit 302 are more detailed than the former. The secondary selector unit 302 is a selector that increases the certainty that what is detected by the primary selector unit 101 of the edge device 1 is the object desired by the user. The detailed selection criteria in the secondary selector unit 302 are, for example, criteria that allow the secondary selector unit 302 to reliably determine that the object in the image is a person when the primary selector unit 101 detects that an object of approximately the same size as a person is captured and that there is a possibility that a person is captured.

[0061] The secondary selector unit 302 recognizes the attributes of the detected object using a recognition learning model M31, for example, to determine whether the detected object is the user's desired object. The recognition learning model M31 is selected from an age recognizer, a glasses-wearing recognizer, a hat-wearing recognizer, a facial orientation recognizer, a posture recognizer, or a detector of people wearing clothes of a specific color (e.g., yellow clothes) provided by the database 310, depending on the object to be detected set for each space, i.e., the user's desired viewing object, and is stored in the storage unit 31 so as to be usable by the processing unit 30. Furthermore, an object recognition process using the recognition learning model M31 is deployed to the cloud server 3, and the processing unit 30 executes the recognition process as the secondary selector unit 302. The secondary selector unit 302 inputs frame images stored by the cloud-side storage unit 301 to the recognition learning model M31, and stores the frame image in the storage unit 31 only if the input frame image shows an object with the desired attributes based on the recognition result output from the recognition learning model M31.

[0062] Instead of or in addition to the recognition learning model M31, a detection model that detects objects with higher accuracy may be used. For example, the primary selector unit 101 may use the detection learning model M11 to detect on the edge device 1 whether an object of approximately the same size as a person is captured, and the secondary selector unit 302 may use a secondary-side detection learning model to detect with high accuracy whether a person is actually captured, using the likelihood that the object is actually a person as a detailed selection criterion. The processing unit 30 may input frame images to the secondary-side learning model, and based on the detection results output from the model, narrow down the frame images to those that are more likely to capture an object and store them in the storage unit 31.

[0063] If text information corresponding to the explanation of the criteria for selecting the frame image is stored in association with the frame image, the secondary selector unit 302 may refer to the text information and determine that the frame image is the viewing object desired by the user. In this case, the text information is used as supplemental information.

[0064] Of the frame images stored by cloud-side storage unit 301, secondary selector unit 302 may delete frame images other than those stored by processing by secondary selector unit 302.

[0065] The selection operation unit 303 controls the operation unit 44 of the client 4 to accept a selection operation of a frame image to be displayed on the display unit 43 of the client 4 connected to the cloud server 3, from among the frame images narrowed down by the secondary selection of the secondary selector unit 302. The processing unit 30 accepts the selection operation on a web page displayed on the client 4, based on a web server module included in the server program P31.

[0066] The image display control unit 304 controls the display unit 43 of the client 4 connected to the cloud server 3 to display the frame images narrowed down by the secondary selection of the secondary selector unit 302. Based on a module for the Web server included in the server program P31, the processing unit 30, as the image display control unit 304, controls the Web page displayed on the client 4 to include the frame images narrowed down by the secondary selection.

[0067] In this way, in image display system 100 that performs each function, edge device 1 performs primary selection by detecting whether or not an object is captured in a frame image captured by camera 2 according to rough selection criteria. Edge device 1 selects frame images in which an object is detected (primary selection) and transmits them to cloud server 3. Cloud server 3 then performs secondary selection by selecting frame images in which the object is captured from the selected frame images according to detailed selection criteria. The selected frame images are stored in storage unit 31 so that they can be viewed by client 4.

[0068] The procedure for each function of the image display system 100 shown in Fig. 5 will be described with reference to a flowchart. Fig. 6 and Fig. 7 are flowcharts showing an example of a processing procedure in the image display system 100.

[0069] The processing unit 10 of the edge device 1 acquires a frame image from the camera 2 using the function of the primary selector unit 101 (step S111). The processing unit 10 inputs the acquired frame image to the detection learning model M11 (step S112) and acquires a detection result (step S113). Based on the detection result, the processing unit 10 determines whether or not there is a possibility that the object desired by the user is captured in the frame image acquired in step S111 (step S114).

[0070] If the processing unit 10 determines that the frame image may contain an object desired by the user (S114: YES), the processing unit 10 stores the frame image in the storage unit 11 in association with time information indicating the time when the frame image was captured (step S115). The primary selector unit 101 of the processing unit 10 ends processing on the acquired frame image. In step S115, the processing unit 10 may also store text information (e.g., "human") indicating the detection target of the detection learning model M11 in association with the frame image in the storage unit 11.

[0071] In step S114, if the processing unit 10 determines that there is no possibility that the object desired by the user is captured in the frame image (S114: NO), the processing unit 10 ends the processing for the acquired frame image. The acquired frame image does not need to be saved in the storage unit 11.

[0072] When a new frame image is stored in the storage unit 11, the processing unit 10 of the edge device 1 transmits the frame image and corresponding time information to the cloud server 3 by the function of the edge device-side transmission unit 102 (step S121). In step S121, the processing unit 10 also transmits identification data of the edge device 1. This allows the cloud server 3 corresponding to multiple spaces to identify in which space the frame image was captured.

[0073] On the cloud server 3 side, the processing unit 30 receives the frame image and time information transmitted from the edge device 1 using the function of the cloud-side storage unit 301 (step S311), stores them in the memory unit 31 (step S312), and ends the receiving and storing process. In step S312, the processing unit 30 stores the identification data of the edge device 1 received together in association with the received data.

[0074] When a new frame image that may possibly show an object is saved in the memory unit 31 by the function of the secondary selector unit 302, the processing unit 30 inputs the new frame image to the recognition learning model M31 (step S321). The processing unit 30 acquires the recognition result (whether or not the attributes are satisfied) from the recognition learning model M31 (step S322), and determines whether or not the object desired by the user is shown in the frame image input to the recognition learning model M31 based on detailed selection criteria (step S323). In step S323, the processing unit 30 may also refer to the text information when text information is stored in association with the frame image.

[0075] If the processing unit 30 determines that the frame image shows the object desired by the user based on the detailed selection criteria (S323: YES), it stores the frame image and the corresponding time information in the storage unit 31 in association with each other (step S324). If the processing unit 30 selects chronologically consecutive frame images in step S324, it may store them as a single video. In step S324, the processing unit 30 may acquire and store a predetermined amount of video before and after the frame image to be stored from the edge device 1. In step S324, the processing unit 30 also stores the correspondingly stored identification data of the edge device 1. The processing unit 30 stores text corresponding to the attributes recognized by the recognition learning model M31 used for the new frame image, i.e., text corresponding to the reason (attribute) for selecting the target frame image, in association with the frame image and the time information (step S325). In step S325, the processing unit 30 stores, for example, text information indicating the attributes of the object recognized by the recognition learning model M31 used in step S321. The process in step S325 means storing text for searching, and is therefore unnecessary if text-based searches are not accepted.

[0076] If the processing unit 30 determines that the frame image does not show the object desired by the user (S323: NO), it deletes the target frame image stored in the memory unit 31 by the cloud-side storage unit 301 from the memory unit 31 (step S326).

[0077] 6 and 7, only the frame images that have been secondary selected by the secondary selector unit 302 are stored in the cloud server 3. In the first embodiment, the frame images and corresponding videos selected by the primary selection in the edge device 1 may be deleted in order from oldest to newest using a FIFO. If the frame images remaining in the cloud server 3 are stored together with text information that explains the attributes of the object that is the basis for their storage, subsequent searches will be easier.

[0078] As described above, by distributing the processing between the edge device 1 and the cloud server 3 for primary selection and secondary selection, it is possible to reduce the processing load on the edge device 1, which has relatively few computing resources. Because the edge device 1 transmits the data for which primary selection has been performed, the amount of communication traffic can be reduced compared to transmitting all captured frame images to the cloud server 3.

[0079] The frame image selected by secondary selection and saved in storage unit 31 of cloud server 3 can be displayed on display unit 43 of client 4. Fig. 8 is a flowchart showing an example of a display control procedure in image display system 100. When a user accesses cloud server 3 using client 4, processing unit 30 of cloud server 3 functions as selection operation unit 303 and image display control unit 304 and starts the following processing.

[0080] The processing unit 30 of the cloud server 3 receives data identifying the target space from the client 4 as the selection operation unit 303 (step S331). In step S331, the processing unit 30 accepts the user's account data, the space's identification data or name, etc. The processing unit 30 may also accept the identification data of the camera 2 included in the image information. In step S331, the processing unit 30 may also accept a selection from a list of the identification data of the edge device 1 that is permitted to access the account data used when the client 4 accessed the cloud server 3, and the corresponding space's identification data or name.

[0081] The processing unit 30 identifies the identification data of the edge device 1 corresponding to the space identified by the received data (step S332), and reads out the frame images associated with the identified identification data of the edge device 1 from the storage unit 31 (step S333). In step S333, the consecutive frame images may be read out as a moving image.

[0082] The processing unit 30, as the selection operation unit 303, receives a search word corresponding to an object (viewing target) desired by the user at the client 4 (step S334). In step S334, the processing unit 30 may cause the client 4 to display a list of thumbnails of the frame images read out at step S333, and may receive a selection of an object desired by the user by a user selection operation of one of the thumbnails.

[0083] The processing unit 30, as the image display control unit 304, extracts frame images associated with text information corresponding to the search word or frame images selected from the list from the frame images read out in step S333 (step S335). The processing unit 30 displays the list of extracted frame images on the display unit 43 of the client 4 (step S336), and ends the process.

[0084] In step S336, if the frame images read out in step S333 are continuous in time series and are a moving image, the processing unit 30 may display a representative frame image as a thumbnail in step S336.

[0085] Of the processing steps shown in FIG. 8, steps S334 and S335 may be omitted, and the list of frame images read out in step S303 may be displayed in step S336.

[0086] 9 and 10 are explanatory diagrams showing an example of image display on the client 4. In FIGS. 9 and 10, an example is shown in which a screen provided by the web server function of the cloud server 3 in response to access from the processing unit 40 of the client 4 is displayed by the web browser function included in the client program P4. A screen 430 shown in FIG. 9 displays a list 431 of spaces that the user operating the client 4 is permitted to view. The list 431 includes an interface 432 with a link to an image display screen that displays images saved in each space. In the example of FIG. 9, a list 431 of "Store (X Chain Store A)," "Store (X Chain Store B)," and "Store (X Chain Store C)" is displayed.

[0087] Fig. 10 is a diagram showing an example of the image display screen 433. The image display screen 433 shown in Fig. 10 is displayed when the user selects the interface 432 for "Store (X Chain Store A)" from the list 431 shown in Fig. 9 using the operation unit 44.

[0088] Image display screen 433 includes area 434 showing a list of frame images that have been selected and saved by primary selector unit 101 and secondary selector unit 302, and area 435 showing a list of object images obtained by extracting areas in which objects appearing in the frame images have been detected. When an object image included in area 435 is selected on image display screen 433, the frame image from which the object image was extracted is highlighted in area 434. By operating operation unit 44 on an object image, a video including the corresponding frame image may be downloaded from data saved on cloud server 3 to storage unit 41 of client 4, or may be temporarily downloaded to a temporary storage medium in processing unit 40. The process of enabling downloading to storage unit 41 of client 4 corresponds to the image saving control unit in the claims.

[0089] The image display system 100 of the first embodiment can accurately select a frame image showing the object desired by the user to view while reducing the processing load on the edge device 1, and display the selected frame image and a predetermined amount of video before and after this frame image on the display unit 43 of the client 4 via the cloud server 3.

[0090] [Second embodiment] In the image display system 100 of the first embodiment, the frame image selected by the edge device 1 is temporarily stored as is in the cloud server 3 by the function of the cloud-side storage unit 301. In the second embodiment, the data of the frame image itself showing the object desired by the user is not stored in the storage unit 31 by the processing of the secondary selector unit 302, and only the time information of the frame image is stored in the storage unit 31.

[0091] Among the configurations of the image display system 100 of the second embodiment, the hardware configuration of each device is the same as that of the image display system 100 of the first embodiment. Therefore, among the configurations of the image display system 100 of the second embodiment below, the same reference numerals are used for the configurations common to the first embodiment, and detailed description thereof will be omitted.

[0092] 11 is a functional block diagram of an image display system 100 according to the second embodiment. In the image display system 100 according to the second embodiment, similarly to the first embodiment, the processing unit 10 of the edge device 1 functions as a primary selector unit 101 and an edge device-side transmitter 102, and the processing unit 30 of the cloud server 3 functions as a cloud-side storage unit 301, a secondary selector unit 302, a selection operation unit 303, and an image display control unit 304.

[0093] In the second embodiment, the processing unit 10 of the edge device 1 also uses the detection learning model M11 as the primary selector unit 101 to select frame images that may contain an object. As in the first embodiment, the detection learning model M11 is selected from the database 310 depending on the object and is made available for use in the detection process. The primary selector unit 101 saves the selected frame images in the storage unit 11 and discards the unselected frame images. The primary selector unit 101 saves a predetermined amount of video before and after the selected frame image. When the primary selector unit 101 selects chronologically consecutive frame images, it may save them as a continuous video.

[0094] In the second embodiment as well, the processing unit 10 of the edge device 1 serves as the edge device-side transmitting unit 102 and transmits the frame image selected by the function of the primary selector unit 101 via the second communication unit 13. In the second embodiment, the edge device-side transmitting unit 102 does not need to transmit a moving image including other frame images before and after the frame image.

[0095] In the second embodiment, the cloud-side storage unit 301 stores the frame images that have undergone primary selection and are transmitted from the edge device 1 in the storage unit 31. Note that the image information stored in the storage unit 31 may be deleted by processing by the secondary selector unit 302, which will be described later.

[0096] In the second embodiment, the secondary selector unit 302, as in the first embodiment, uses the recognition learning model M31 to select a frame image showing an object desired by the user from the frame images stored in the storage unit 31. As in the first embodiment, the recognition learning model M31 is selected from the database 310 according to the object and made available for use in the recognition process. The secondary selector unit 302 associates time information corresponding to the selected frame image with the identification data of the edge device 1 that sent the image, the identification data of the space, or the identification data of the camera 2, and stores the time information in the storage unit 31. The secondary selector unit 302 may delete frame images that have undergone secondary selection processing from the storage unit 31. This prevents frame images captured at various locations from being stored and accumulated in the cloud server 3 accessible from each client 4.

[0097] In the second embodiment, the selection operation unit 303 also controls the operation unit 44 of the client 4 to accept the user's selection operation of image information to be displayed on the display unit 43 of the client 4 connected to the cloud server 3 from the image information narrowed down by the secondary selection of the secondary selector unit 302.

[0098] In the second embodiment, the image display control unit 304 transmits a video including a frame image corresponding to the stored time information selected by the secondary selection of the secondary selector unit 302 from the edge device 1 to the client 4 and displays it on the display unit 43. The image display control unit 304 may temporarily acquire a frame image corresponding to the time information specified by the client 4 and a video including the frame image, and transmit the acquired video to the client 4. The image display control unit 304 may redirect the transmission of the video hosted by the edge device 1 to the client 4 while temporarily acquiring the frame image corresponding to the time information specified by the client 4. When the edge device 1 transmits the video directly, it is possible to prevent data from being stored in the cloud server 3. In this case, the first image display program P1 of the edge device 1 includes a module that functions as a host for video distribution. The edge device 1 may also store unique identification data (e.g., MAC addresses) of clients 4 that are permitted to access the host as a whitelist, allowing only permitted clients 4 to stream or download the video. The process of allowing downloading to the storage unit 41 of the client 4 corresponds to the image storage control unit in the claims.

[0099] 12 and 13 are flowcharts showing an example of a processing procedure in the image display system 100 of the second embodiment. Among the processing procedures shown in Fig. 12 and 13, steps common to the processing procedures shown in Fig. 6 and Fig. 7 of the first embodiment are assigned the same step numbers, and detailed explanations thereof will be omitted.

[0100] The processing procedures of the primary selector unit 101, edge device-side transmission unit 102, and cloud-side storage unit 301 are the same as those shown in Figures 6 and 7 of the first embodiment. In the second embodiment, if the processing unit 30 of the cloud server 3 determines in step S323 using the function of the secondary selector unit 302 that the frame image shows an object desired by the user based on the detailed selection criteria (S323: YES), it stores time information corresponding to the frame image in the storage unit 31 (step S334). In step S334, the processing unit 30 stores identification data of the edge device 1 that transmitted the frame image in association with the time information. The processing unit 30 may delete the stored new frame image from the storage unit 31 (step S335).

[0101] The processing unit 30 stores text information corresponding to the attributes recognized by the recognition learning model M31 used for the new frame image in step S321, i.e., text information corresponding to the basis (attributes) for selecting the target frame image, in association with time information (step S336). The processing of step S336 means storing text information for search purposes, and is therefore not necessary if text-based searches are not accepted.

[0102] If the processing unit 30 determines that the frame image does not show the object desired by the user (S323: NO), it deletes the target frame image stored in the memory unit 31 by the cloud-side storage unit 301 from the memory unit 31 (S326) and terminates the processing.

[0103] Fig. 14 is a flowchart showing an example of a processing procedure by the selection operation unit 303 and the image display control unit 304 of the second embodiment. Of the processing procedures shown in Fig. 14, steps common to the processing procedures shown in Fig. 8 of the first embodiment are assigned the same step numbers, and detailed descriptions thereof will be omitted.

[0104] In the second embodiment, the processing unit 30 of the cloud server 3, as the image display control unit 304, identifies the identification data of the edge device 1 corresponding to the target space based on the data received from the client 4 (S332), and reads out the time information associated with the identification data of the identified edge device 1 from the storage unit 31 (step S341). The processing unit 30 requests the edge device 1 of the identified identification data to transmit a video including a frame image corresponding to the read time information (a video for a predetermined time before and after this frame image) from the edge device 1 (step S342). The processing unit 30 acquires data of the transmission host of the video from the edge device 1 that is the request destination (step S343), redirects the data to the client 4 that is the access source (step S344), displays the video on the client 4 (step S345), and ends the process.

[0105] In the image display system 100 of the second embodiment, the screen 430 and image display screen 433 as shown in Figs. 9 and 10 of the first embodiment can be displayed on the client 4, and a list of frame images showing the viewing target desired by the user can be displayed so that the user can select it. It may also be possible to stream or download image information such as a moving image corresponding to the selected frame image from the edge device 1 to the client 4. Here, the process of enabling downloading to the storage unit 41 of the client 4 corresponds to the image saving control unit in the claims.

[0106] As shown in the second embodiment, by saving as few frame images and videos as possible in the cloud server 3, it becomes possible to display necessary images without saving images of subjects, particularly customers, photographed in each space in the cloud server 3. This makes it possible to view frame images or videos that are determined to show a specific object with a higher degree of accuracy, while taking personal information into consideration.

[0107] [Third embodiment] In the image display system 100 of the first embodiment, the frame image selected by the edge device 1 is temporarily stored as is in the cloud server 3 by the function of the cloud-side storage unit 301. In the third embodiment, the processing on the cloud server 3 side and the storage unit 31 are distributed among multiple server computers. Also in the third embodiment, as in the second embodiment, the data that is left in the storage unit 31 of the cloud server 3 is not the data itself of the frame image in which the user's desired object is displayed, but only the time information of the frame image.

[0108] Among the configurations of the image display system 100 of the third embodiment, the hardware configuration of each device is the same as that of the image display system 100 of the first embodiment. Therefore, among the configurations of the image display system 100 of the third embodiment below, the same reference numerals are used for the configurations common to the first embodiment, and detailed description thereof will be omitted.

[0109] 15 is a functional block diagram of an image display system 100 according to the third embodiment. In the image display system 100 according to the third embodiment, the cloud server 3 is made up of two servers configured with different hardware.

[0110] In the image display system 100 of the third embodiment, the processing unit 10 of the edge device 1 also functions as a primary selector unit 101 and an edge device side transmission unit 102, the processing unit 30 of the first cloud server 3 functions as a cloud side storage unit 301 and a secondary selector unit 302, and the processing unit 30 of the second cloud server 3 functions as an image display control unit 304.

[0111] The functions of the edge device 1 in the third embodiment are similar to the functions of the edge device 1 in the first and second embodiments.

[0112] In the third embodiment, the cloud-side storage unit 301 functioning under the processing unit 30 of the first cloud server 3 stores the frame images that have been sent from the edge device 1 and that have undergone the primary selection in the storage unit 31 .

[0113] Furthermore, a secondary selector unit 302, which is functioned by the processing unit 30 of the first cloud server 3, uses a recognition learning model M31 to select a frame image showing an object desired by the user from the frame images stored in the memory unit 31 of the first cloud server 3. As in the first embodiment, the recognition learning model M31 is selected from the database 310 and made available for use in conjunction with the recognition process. The secondary selector unit 302 associates time information corresponding to the selected frame image with identification data of the edge device 1 that sent the image, identification data of the space where the image was taken, or identification data of the camera 2, and stores the time information in the memory unit 31 of the second cloud server 3.

[0114] The image display control unit 304, which is functioning with the processing unit 30 of the second cloud server 3, transmits a video including a frame image corresponding to the stored time information selected by the secondary selection of the secondary selector unit 302 from the edge device 1 to the client 4 and displays it on the display unit 43. The image display control unit 304 may temporarily acquire a frame image corresponding to the time information specified by the client 4 and a video including the frame image and transmit them to the client 4. The image display control unit 304 may redirect the transmission of the video hosted by the edge device 1 to the client 4 while temporarily acquiring the frame image corresponding to the time information specified by the client 4. In the third embodiment, too, when a video is transmitted directly from the edge device 1, it is possible to prevent data from being stored in the cloud server 3. In this case, the first image display program P1 of the edge device 1 of the third embodiment includes a module that functions as a host for video distribution. The edge device 1 may also store unique identification data (e.g., MAC addresses) of clients 4 that are permitted to access the host as a whitelist, allowing only permitted clients 4 to stream or download the video. The process of making it possible to download the image data to the storage unit 41 of the client 4 corresponds to the image storage control unit in the claims.

[0115] 15, in the third embodiment, the cloud server is divided into a first cloud server 3 and a second cloud server 3. Frame images are sent to and stored in the first cloud server 3 in a secure environment, but are not sent to the storage unit 31 of the second cloud server 3 that performs the functions of the selection operation unit 303 and the image display control unit 304, i.e., the cloud server 3 in an open environment that accepts access from clients 4. This makes it possible to avoid a situation in which frame images taken at various locations are stored and accumulated in the cloud server 3 in an open environment that is accessible from each client 4.

[0116] The processing procedures of each function of the image display system 100 of the third embodiment are the same as those of the second embodiment, except that the cloud server 3 is composed of two servers. The processing unit 30 of the first cloud server 3 executes the processing (S311, S312) as the cloud-side storage unit 301 and the processing (S321-S323, S326, S334-S336) as the secondary selector unit 302 in the processing procedures shown in Figures 12 and 13 of the second embodiment. Furthermore, the processing unit 30 of the second cloud server 3 executes the processing procedures shown in Figure 14 of the first embodiment as the selection operation unit 303 and the image display control unit 304.

[0117] In the image display system 100 of the third embodiment, the screen 430 and the image display screen 433 as shown in Figs. 9 and 10 of the first embodiment can be displayed on the client 4, and a list of frame images showing the viewing target desired by the user can be displayed so that the user can select it. It may also be possible to stream or download a video corresponding to the selected frame image from the edge device 1 to the client 4. Here, the process of enabling downloading to the storage unit 41 of the client 4 corresponds to the image saving control unit in the claims.

[0118] As shown in the third embodiment, the second cloud server 3 accessible from the client 4 does not store frame images and videos, but stores only time information, so that necessary images can be displayed without storing images of objects, particularly customers, photographed in each space in the cloud server 3 accessible from the client 4. This allows frame images or videos that are determined to show a specific object with a higher degree of accuracy to be viewable, while taking personal information into consideration.

[0119] [Fourth embodiment] In the image display system 100 of the first embodiment, the frame image selected by the edge device 1 is temporarily stored as is in the cloud server 3 by the function of the cloud-side storage unit 301. In the fourth embodiment, the frame image is not transmitted from the edge device 1 to the cloud server 3, thereby further reducing the communication load. The cloud server 3 of the fourth embodiment leaves only the time information corresponding to the frame image showing the object desired by the user in the storage unit 31 of the cloud server 3.

[0120] Among the configurations of the image display system 100 of the fourth embodiment, the hardware configuration of each device is the same as that of the image display system 100 of the first embodiment. Therefore, among the configurations of the image display system 100 of the fourth embodiment below, the same reference numerals are used for the configurations common to the first embodiment, and detailed description thereof will be omitted.

[0121] 16 is a functional block diagram of an image display system 100 according to the fourth embodiment. In the image display system 100 according to the fourth embodiment, similarly to the first embodiment, the processing unit 10 of the edge device 1 functions as a primary selector unit 101 and an edge device-side transmitter 102, and the processing unit 30 of the cloud server 3 functions as a cloud-side storage unit 301, a secondary selector unit 302, and an image display control unit 304.

[0122] In the fourth embodiment, the processing unit 10 of the edge device 1 functions as the primary selector unit 101 and receives frame images captured by the camera 2 via the first communication unit 12. For the acquired frame images, the primary selector unit 101 obtains text information representing sentences or words describing an object or person appearing in the frame image using a language model M12. The language model M12 is multimodal and capable of receiving images as input, or a model called a VLM (Visual Language Model) is selected from the database 310 and stored in the storage unit 11. Furthermore, a languageization process using the language model M12 is deployed in the edge device 1. The primary selector unit 101 instructs the language model M12 to output sentences or words describing the person or object appearing in the frame image, and obtains the sentences or words output from the language model M12. The processing unit 10 functions as the primary selector unit 101 and selects the target frame image if it is determined from the obtained sentences or words that the object desired by the user is likely to be captured. The primary selector unit 101 may store the selected frame image in the storage unit 11, and discard the unselected frame images. The primary selector unit 101 stores a predetermined time of video before and after the selected frame image in the storage unit 11. When the primary selector unit 101 selects chronologically consecutive frame images, it may store them as a continuous video. When the processing unit 10, functioning as the primary selector unit 101, cannot detect an object such as a person, it ends processing for the target frame image, does not perform any further processing, and performs processing for the next frame image.

[0123] In the fourth embodiment, the edge device-side transmitting unit 102 transmits, via the second communication unit 13, text information, which is image information corresponding to the frame image selected by the primary selector unit 101, and time information corresponding to the frame image, to the cloud server 3. At this time, the edge device-side transmitting unit 102 transmits the text information and time information in association with the identification data of the edge device 1. As in the first embodiment, the time information is information indicating the time corresponding to the frame image, a time stamp, or an elapsed time or count from a specific time.

[0124] In the fourth embodiment, the cloud-side storage unit 301 receives text information describing the selected frame image, time information, and identification data of the edge device 1 transmitted from the edge device 1 by the edge device-side transmission unit 102, and stores them in the memory unit 31. Image information does not need to be transmitted from the edge device 1 to the cloud server 3.

[0125] In the fourth embodiment, the secondary selector unit 302 selects a frame image showing an object desired by the user according to detailed selection criteria, using text information stored in the memory unit 31 by the cloud-side storage unit 301. The detailed selection criteria are criteria for more reliably detecting whether an object desired by the user is shown, similar to the selection criteria determined by the secondary selector unit 302 in the first embodiment.

[0126] The secondary selector unit 302 uses the language model M32 for sentences or words describing what is shown in the frame image associated with the target time information and stored in the storage unit 31, selects time information indicating the time when the frame image showing the object was captured, and stores the selected time information together with the corresponding text information in the storage unit 31. The secondary selector unit 302 may perform selection by processing whether or not a specific word is included, without using the language model M32. Because the selection criteria by the secondary selector unit 302 are more detailed (stricter) than the selection criteria by the primary selector unit 101, the text information and corresponding time information stored in the storage unit 31 have a smaller data volume than the received text information and corresponding time information. The secondary selector unit 302 may delete text information stored by the cloud-side storage unit 301 other than the text information stored by the processing of the secondary selector unit 302.

[0127] 17 and 18 are flowcharts showing an example of a processing procedure in the image display system 100 of the fourth embodiment. Among the processing procedures shown in Fig. 17 and Fig. 18, steps common to the processing procedures shown in Fig. 6 and Fig. 7 of the first embodiment are assigned the same step numbers, and detailed explanations thereof will be omitted.

[0128] The primary selector unit 101 inputs the frame image acquired in step S111 together with the instruction sentence into the language model M12 (step S116), and acquires text information such as sentences or words output from the language model M (step S117). The processing unit 10 uses the acquired text information to determine whether or not there is a possibility that an object desired by the user is captured in the frame image acquired in step S111 (step S118).

[0129] If the processing unit 10 determines that there is a possibility that the object desired by the user is captured in the frame image (S118: YES), it stores the text information acquired in step S117 and time information indicating the time when the frame image was captured in the storage unit 11 (step S119). The primary selector unit 101 of the processing unit 10 ends the processing for the acquired frame image.

[0130] In step S118, if the processing unit 10 determines from the text information acquired in step S117 that there is no possibility that the object desired by the user is shown (S118: NO), the processing unit 10 ends the processing for the acquired frame image. In this case, the acquired text information does not need to be saved in the storage unit 11.

[0131] When new text information is stored in the storage unit 11, the processing unit 10 of the edge device 1 transmits the text information and corresponding time information to the cloud server 3 by the function of the edge device-side transmitting unit 102 (step S122). In step S122, the processing unit 10 transmits the identification data of the edge device 1 together with the text information.

[0132] On the cloud server 3 side, the processing unit 30 receives the text information and time information transmitted from the edge device 1 using the function of the cloud-side storage unit 301 (step S313), stores them in the memory unit 31 (step S314), and ends the receiving and storing process. In step S314, the processing unit 30 stores the identification data of the edge device 1 received together in association with the received information.

[0133] When new text information is stored in the storage unit 31 by the function of the secondary selector unit 302, the processing unit 30 inputs the new text information to the language model M32 (step S351). The processing unit 30 acquires text information corresponding to the description for the frame image from the language model M32 (step S352). Based on the acquired text information, the processing unit 30 determines whether or not the corresponding frame image shows an object desired by the user in accordance with detailed selection criteria (step S353).

[0134] If the processing unit 30 determines, according to the detailed (strict) selection criteria, that the frame image corresponding to the text information shows an object desired by the user (S353: YES), it stores time information indicating the time when the frame image was captured in the storage unit 31 (step S354). In step S354, the processing unit 30 also stores, in association with the time information, the identification data of the edge device 1 that was stored in association with the time information in step S314. For the new text information input to the language model M32 in step S351, the processing unit 30 stores, in association with the time information, text information corresponding to the attribute recognized as the desired object (step S355). The processing of step S355 means storing text information for search purposes, and is therefore unnecessary if a text-based search is not accepted.

[0135] If the processing unit 30 determines that the frame image corresponding to the text information does not show the object desired by the user (S353: NO), it may delete the target text information stored in the memory unit 31 by the cloud-side storage unit 301 from the memory unit 31 (step S356).

[0136] 17 and 18, time information corresponding to the frame images that have been secondary selected by the secondary selector unit 302 is stored in the cloud server 3. If text information corresponding to the attributes of the object that is the basis for storing the time information is stored in the cloud server 3 in association with the time information, it will be easier to search for it later.

[0137] As described above, by distributing the processing between the edge device 1 and the cloud server 3 for the primary selection and the secondary selection, it is possible to reduce the processing load on the edge device 1. Text information corresponding to the description of the frame image for which the edge device 1 has performed the primary selection is sent to the cloud server 3, which allows for a significant reduction in the amount of communication traffic.

[0138] In the fourth embodiment as well, text information selected by secondary selection and stored in the storage unit 31 of the cloud server 3 can be selectably displayed on the display unit 43 of the client 4. Fig. 19 is a flowchart showing an example of the processing procedure by the selection operation unit 303 and the image display control unit 304 of the fourth embodiment. When a user accesses the cloud server 3 using the client 4, the processing unit 30 of the cloud server 3 starts the following processing as the image display control unit 304. Of the processing procedures shown in Fig. 19, steps that are common to the processing procedures shown in Fig. 8 of the first embodiment are assigned the same step numbers, and detailed descriptions thereof will be omitted.

[0139] When the processing unit 30 of the cloud server 3 receives data specifying the target space of the client 4 as the selection operation unit 303 (S331), it specifies the identification data of the edge device 1 corresponding to the space specified by the received data (S332). Time information associated with the identification data of the specified edge device 1 is read from the storage unit 31 (step S363). In step S363, the continuous time information may be read as time information corresponding to continuous videos.

[0140] The processing unit 30 receives a search word corresponding to an object desired by the user from the client 4 via the selection operation unit 303 (step S364). The processing unit 30 extracts time information associated with text information including text corresponding to the received search word from the time information read out in step S363 (step S365).

[0141] The processing unit 30 acquires a frame image corresponding to the extracted time information from the storage unit 11 of the edge device 1 (step S366). The processing unit 30, as the image display control unit 304, displays a list of thumbnails using the acquired frame images on the client 4 (step S367). The processing unit 30 accepts the selection of any frame image from the thumbnail list via the selection operation unit 303 (step S368).

[0142] The processing unit 30 requests the edge device 1 that is the sender to send a video including the selected frame image from the edge device 1 (step S369). The processing unit 30, as the image display control unit 304, acquires data of the sending host of the video from the edge device 1 that is the request destination (step S370), redirects the data to the client 4 that is the access source (step S371), displays the video on the client 4 (step S372), and ends the processing.

[0143] The image display system 100 of the fourth embodiment may have a configuration in which the cloud server 3 is divided into two, similar to the third embodiment.

[0144] In the image display system 100 of the fourth embodiment, the screen 430 and image display screen 433 as shown in Figs. 9 and 10 of the first embodiment can be displayed on the client 4, and a list of frame images showing the viewing target desired by the user can be displayed so that the user can select it. It may also be possible to stream or download image information such as a moving image corresponding to the selected frame image from the edge device 1 to the client 4. Here, the process of enabling downloading to the storage unit 41 of the client 4 corresponds to the image saving control unit in the claims.

[0145] In the fourth embodiment, the cloud server 3 receives and stores, by the function of the cloud-side storage unit 301 of the cloud server 3, not the frame image, but text information corresponding to a description of the frame image in which an object may appear, and time information about the time when the frame image in which the object may appear was captured. The frame image captured by the edge device 1 is not transmitted to the cloud server 3, even temporarily. The size of the text information, such as sentences or words describing what is being displayed, is overwhelmingly smaller than that of the frame image, making it possible to reduce the amount of data communication transmitted from the edge device 1 to the cloud server 3. Furthermore, because the frame image and the video including the frame image are not transmitted from the edge device 1 to the cloud server 3, an image display system 100 can be provided that takes personal information into consideration.

[0146] The embodiments disclosed above are illustrative in all respects and are not restrictive. The scope of the present invention is defined by the claims, and includes all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0147] 1. Edge Devices 10 Processing section 101 Primary selector section 102 Edge device side transmitter 11 Storage section P1 First image display program 3. Cloud Server 30 Processing section 301 Cloud storage unit 302 Secondary Selector 304 Image display control unit 31 Storage section 310 Database P32 Second image display program

Claims

1. a primary selector unit that performs a primary selection of frame images on the edge device side, which is a process of selecting frame images that may include a user's desired viewing target by detecting the user's desired viewing target from frame images input from the camera according to rough selection criteria; and an edge device-side transmitter that transmits image information corresponding to the frame image selected by the primary selector from the edge device side to a cloud side; a cloud-side storage unit at the cloud side that stores the image information transmitted by the edge device-side transmission unit; a secondary selector unit on the cloud side that performs a secondary selection of frame images, which is a process of selecting frame images showing a user-desired viewing subject in accordance with detailed selection criteria using image information stored in the cloud-side storage unit; An image display system comprising: an image display control unit on the cloud side that uses time information corresponding to the frame image selected by secondary selection in the secondary selector unit to read a video showing the user's desired viewing subject from a storage device on the edge device side and control the video to be displayed on a display unit on the cloud side.

2. 2. The image display system according to claim 1, wherein the primary selector unit outputs text information relating to a frame image in which the viewing subject may appear as a result of the primary selection.

3. The image display system according to claim 2, characterized in that the edge device side transmission unit transmits text information regarding a frame image in which the object to be viewed may appear and time information when the frame image in which the object to be viewed may appear was taken to the cloud side.

4. The image display system described in claim 2 or claim 3, characterized in that the image information transmitted by the edge device side transmission unit and stored in the cloud side storage unit is the text information output by the primary selector unit, and the secondary selector unit uses the text information to perform secondary selection of the frame images.

5. a selection operation unit for allowing a user to select a desired image from the images displayed on the cloud-side display unit; The image display system according to claim 1, further comprising an image saving control unit that controls the reading of a video including an image selected by a user using the selection operation unit from a storage device on the edge device side and saving the video in a storage device of a device other than the edge device.

6. 4. The image display system according to claim 1, wherein the secondary selector unit performs secondary selection of the frame images using a VLM (Visual Language Model) or an LLM (Large Language Model).

7. Computer, a primary selector unit that performs a primary selection of frame images, which is a process of detecting a user's desired viewing target from frame images input from a camera on the edge device side according to rough selection criteria, and selecting frame images that may contain the user's desired viewing target; an edge device-side transmitter that transmits image information corresponding to the frame image selected by the primary selector from the edge device side to a cloud side; a cloud-side storage unit at the cloud side that stores the image information transmitted by the edge device-side transmission unit; a secondary selector unit on the cloud side that performs a secondary selection of frame images, which is a process of selecting frame images showing a user-desired viewing subject in accordance with detailed selection criteria using image information stored in the cloud-side storage unit; An image display program for functioning as an image display control unit on the cloud side that uses time information corresponding to the frame image selected by secondary selection in the secondary selector unit to read a video showing the user's desired viewing subject from a storage device on the edge device side and control it to be displayed on a display unit on the cloud side.

Citation Information

Patent Citations

  • Video monitoring system

    JP2009060261A

  • Detection recognizing system

    JP2018088157A

  • Retrieval program, retrieval method and information processing apparatus operating retrieval program

    JP2019045894A

  • On-vehicle sensing device and sensor parameter optimization device

    JP2021144689A

  • Information processing system, program and the like

    JP2023126102A