Program, method, information processing device, and system
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-04-08
AI Technical Summary
Existing surveillance camera systems struggle with handling multiple types of tracking targets in a single frame, leading to inconvenience during video editing.
A system that utilizes multiple pre-stored analytical models to detect objects, presents detection markers indicating detection timings, and generates edited videos based on user instructions, allowing for intuitive video editing.
Improves convenience in editing videos by enabling users to easily manage and edit multiple objects within a frame, ensuring no important events are missed.
Smart Images

Figure 00000000_0001_ABST 
Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, a method, an information processing device, and a system. [Background technology]
[0002] Patent Document 1 describes a technology applicable to a surveillance camera system for specifying a tracking target before or during tracking of an object. In Patent Document 1, a user selects an object (tracking target) that they wish to enlarge from among extracted tracking target candidates, and obtains a desired enlarged image (zoomed image). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-251940 Summary of the Invention [Problem to be solved by the invention]
[0004] In Patent Document 1, one type of tracking target is extracted in one frame. However, one frame may contain multiple types of tracking targets. When editing videos based on images captured by a surveillance camera or the like, there is a demand for improved convenience.
[0005] An object of the present disclosure is to improve convenience when editing videos based on images captured by a surveillance camera or the like. [Means for solving the problem]
[0006] A program to be executed by a computer having a processor and a memory causes the processor to execute the following steps: acquiring a captured video, analyzing the acquired video using a plurality of pre-stored analytical models, each of which has been trained to detect an object, presenting to a user detection markers that indicate the timing at which the object is detected by the analytical models and that are associated with a timeline associated with the video, and generating an edited video from the captured video in response to a user instruction based on the detection markers. [Effects of the Invention]
[0007] According to the present disclosure, convenience can be improved when editing videos based on images captured by a surveillance camera or the like. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of a system. [Figure 2] FIG. 2 is a block diagram illustrating a configuration of a terminal device shown in FIG. [Figure 3] FIG. 2 is a diagram illustrating a functional configuration of a server. [Figure 4] FIG. 2 is a diagram showing the data structure of a photographing device DB. [Figure 5] FIG. 10 is a diagram illustrating the data structure of a detection log DB. [Figure 6] 10 is a flowchart showing an operation when information detected by image analysis is recorded in a log. [Figure 7] FIG. 10 is a schematic diagram illustrating an editing screen displayed on a display of a terminal device. [Figure 8] FIG. 2 is a schematic diagram illustrating a part of an editing screen displayed on a display of a terminal device. [Figure 9] 10A and 10B are schematic diagrams illustrating other examples of the editing screen displayed on the display of the terminal device. [Figure 10] FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the following description, the same components are denoted by the same reference numerals. The names and functions of the components are also the same. Therefore, detailed descriptions thereof will not be repeated.
[0010] <Summary> The system according to this embodiment analyzes captured video using multiple image analysis models, presents objects that allow users to understand which model detected the target object, and creates an edited video based on the object designation.
[0011] <1 Overall system configuration> Fig. 1 is a block diagram showing an example of the overall configuration of a system 1. The system 1 shown in Fig. 1 includes, for example, a terminal device 10, a server 20, an image capturing device 31, a sensor 32, and a speaker 33. The terminal device 10, the server 20, and the speaker 33 are communicatively connected via, for example, a network 80.
[0012] 1 shows an example in which the system 1 includes two terminal devices 10, but the number of terminal devices 10 included in the system 1 is not limited to two. The terminal devices 10 are terminals carried by camera users. The number of terminal devices 10 included in the system 1 may be less than three, or may be three or more.
[0013] In this embodiment, a collection of multiple devices may be considered as one server. The allocation of multiple functions required to realize the server 20 according to this embodiment to one or more pieces of hardware can be determined appropriately in consideration of the processing capacity of each piece of hardware and / or the specifications required for the server 20.
[0014] 1 may be, for example, a mobile terminal such as a smartphone or a tablet, a desktop personal computer (PC), a laptop PC, or a wearable terminal such as a head mounted display (HMD) or a wristwatch terminal.
[0015] The terminal device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage 16, and a processor 19.
[0016] The communication IF 12 is an interface for inputting and outputting signals so that the terminal device 10 can communicate with devices in the system 1, such as the server 20, for example.
[0017] The input device 13 is a device for receiving input operations from a user (for example, a touch panel, a touch pad, a pointing device such as a mouse, a keyboard, etc.).
[0018] The output device 14 is a device (such as a display or speaker) for presenting information to the user.
[0019] The memory 15 is for temporarily storing programs and data to be processed by the programs, and is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0020] The storage 16 is for storing data, and is, for example, a flash memory or a hard disk drive (HDD).
[0021] The processor 19 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, and the like.
[0022] The server 20 is realized by, for example, a computer connected to a network 80. As shown in FIG.
[0023] The communication IF 22 is an interface for inputting and outputting signals so that the server 20 can communicate with devices in the system 1, such as the terminal device 10, for example.
[0024] The input / output IF 23 functions as an interface with an input device for receiving input operations from the user and an output device for presenting information to the user.
[0025] The memory 25 is for temporarily storing programs and data to be processed by the programs, and is a volatile memory such as a DRAM.
[0026] The storage 26 is for storing data, and is, for example, a flash memory or a HDD.
[0027] The processor 29 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0028] The image capturing device 31 is a device that receives light using a light receiving element and outputs the light as an image signal. The image capturing device 31 captures images at a frame rate that allows a series of movements to be recognized as a video. A frame rate that allows a series of movements to be recognized as a video is, for example, about 30 fps. The image capturing device 31 may be, for example, a camera that can capture a wide range of images in a 360-degree circle. In this case, the image capturing device 31 is realized by, for example, a camera with an ultra-wide-angle lens or a fisheye lens.
[0029] The image capturing device 31 is installed at a position where it can see the entire site without any obstructions. If one image capturing device 31 cannot capture the entire site, multiple image capturing devices 31 are installed. When multiple image capturing devices 31 are installed, for example, a section of the site to be captured by each image capturing device 31 is set in advance. The image capturing device 31 outputs the acquired image signal to the server 20.
[0030] <1.1 Terminal device configuration> Fig. 2 is a block diagram showing an example configuration of the terminal device 10 shown in Fig. 1. The terminal device 10 shown in Fig. 2 is realized by a mobile terminal, a PC, or a wearable terminal. As shown in Fig. 2, the terminal device 10 includes a communication unit 120, an input device 13, an output device 14, an audio processing unit 17, a microphone 171, a speaker 172, a camera 161, a position information sensor 150, a storage unit 180, and a control unit 190. The blocks included in the terminal device 10 are electrically connected by, for example, a bus or the like.
[0031] The communication unit 120 performs processing such as modulation and demodulation for the terminal device 10 to communicate with other devices. The communication unit 120 performs transmission processing on the signal generated by the control unit 190 and transmits it to the outside (for example, the server 20). The communication unit 120 performs reception processing on the signal received from the outside and outputs it to the control unit 190.
[0032] The input device 13 is a device for inputting instructions or information by a user operating the terminal device 10. The input device 13 is realized, for example, by a touch-sensitive device 131 or the like, which inputs instructions by touching an operation surface. When the terminal device 10 is a PC or the like, the input device 13 may be realized by a reader, keyboard, mouse, or the like. The input device 13 converts instructions input by the user into electrical signals and outputs the electrical signals to the control unit 190. The input device 13 may include, for example, a receiving port that receives electrical signals input from an external input device.
[0033] The output device 14 is a device for presenting information to a user operating the terminal device 10. The output device 14 is realized, for example, by a display 141 or the like. The display 141 displays data according to the control of the control unit 190. The display 141 is realized, for example, by an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display or the like.
[0034] The audio processing unit 17 performs, for example, digital-to-analog conversion processing of an audio signal. The audio processing unit 17 converts a signal provided from the microphone 171 into a digital signal and provides the converted signal to the control unit 190. The audio processing unit 17 also provides the audio signal to the speaker 172. The audio processing unit 17 is realized, for example, by a processor for audio processing. The microphone 171 receives audio input and provides an audio signal corresponding to the audio input to the audio processing unit 17. The speaker 172 converts the audio signal provided from the audio processing unit 17 into audio and outputs the audio to the outside of the terminal device 10.
[0035] The camera 161 is a device that receives light with a light receiving element and outputs the light as an image signal.
[0036] The position information sensor 150 is a sensor that detects the position of the terminal device 10, and is, for example, a GPS (Global Positioning System) module. The GPS module is a receiving device used in a satellite positioning system. In the satellite positioning system, signals are received from at least three or four satellites, and the current position of the terminal device 10 equipped with the GPS module is detected based on the received signals. The position information sensor 150 may detect the current position of the terminal device 10 from the position of the wireless base station to which the terminal device 10 is connected.
[0037] The storage unit 180 is realized by, for example, the memory 15, the storage 16, etc., and stores data and programs used by the terminal device 10. The storage unit 180 stores, for example, user information 181.
[0038] The user information 181 includes, for example, information about the user who uses the terminal device 10. The information about the user includes, for example, information for identifying the user, the user's name, age, address, date of birth, telephone number, email address, etc.
[0039] The control unit 190 is realized by the processor 19 reading a program stored in the storage unit 180 and executing instructions included in the program. The control unit 190 controls the operation of the terminal device 10. The control unit 190 functions as an operation reception unit 191, a transmission / reception unit 192, and a presentation control unit 193 by operating in accordance with the program.
[0040] Operation reception unit 191 performs processing for receiving instructions or information input from input device 13. Specifically, for example, operation reception unit 191 receives information based on instructions input from touch-sensitive device 131 or the like. The instructions input via touch-sensitive device 131 or the like are, for example, editing instructions from the user.
[0041] Furthermore, the operation reception unit 191 receives voice instructions input from the microphone 171. Specifically, for example, the operation reception unit 191 receives a voice signal that is input from the microphone 171 and converted into a digital signal by the voice processing unit 17. For example, the operation reception unit 191 analyzes the received voice signal and extracts a predetermined noun, thereby acquiring an instruction from the user.
[0042] The transmitting / receiving unit 192 performs processing for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 in accordance with a communication protocol. Specifically, for example, the transmitting / receiving unit 192 transmits an editing instruction input by a user to the server 20. In addition, the transmitting / receiving unit 192 receives information about the user from the server 20.
[0043] The presentation control unit 193 controls the output device 14 to present information provided from the server 20 to the user. Specifically, for example, the presentation control unit 193 causes the information transmitted from the server 20 to be displayed on the display 141. In addition, the presentation control unit 193 causes the information transmitted from the server 20 to be output from the speaker 172.
[0044] <1.2 Functional configuration of the server> 3 is a diagram showing an example of the functional configuration of the server 20. As shown in FIG. 3, the server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203.
[0045] The communication unit 201 performs processing for the server 20 to communicate with external devices.
[0046] The storage unit 202 includes, for example, an imaging device database (DB) 2021, a detection log database (DB) 2022, and the like.
[0047] The photographing device DB 2021 is a database for storing information about cameras installed for photographing, as will be described in detail later.
[0048] The detection log DB 2022 is a database for storing information about objects detected by image analysis, etc. Details will be described later.
[0049] The first trained model 2023 is a model generated by having a machine learning model perform machine learning in accordance with a model learning program. The first trained model 2023 is, for example, a parameterized composite function formed by combining multiple functions that performs predetermined inference based on input data. The parameterized composite function is defined by a combination of multiple adjustable functions and parameters. The trained model according to this embodiment may be any parameterized composite function that meets the above requirements. The same applies to the second trained model 2024 and the third trained model 2025.
[0050] For example, if the first trained model 2023 is generated using a forward propagation type multi-layer network, the parameterized composite function is defined as a combination of, for example, a linear relationship between each layer using a weight matrix, a non-linear relationship (or a linear relationship) using an activation function in each layer, and a bias. The weight matrix and bias are called parameters of the multi-layer network. The form of the parameterized composite function as a function changes depending on how the parameters are selected. In a multi-layer network, by appropriately setting the constituent parameters, it is possible to define a function that can output a desirable result from the output layer. The same applies to the second trained model 2024 and the third trained model 2025.
[0051] The multi-layered network according to the present embodiment may be, for example, a deep neural network (DNN), which is a multi-layered neural network that is the subject of deep learning. As the DNN, for example, a convolution neural network (CNN) that targets images may be used.
[0052] The first trained model 2023 is a model that learns to output information about a human based on an input image when the image is input. The first trained model 2023 may be, for example, different trained models that are trained with different learning data depending on the information to be output.
[0053] For example, the first trained model 2023 is trained to output information about human attributes such as gender, age, race, and occupational status when an image is input. Specifically, in this embodiment, the first trained model 2023 identifies human attributes based on information about the physique, face, hairstyle, clothing, and belongings of the human included in the image. Furthermore, for example, the first trained model 2023 may be trained to output information about a specific person when an image is input. Specifically, in this embodiment, the first trained model 2023 may identify a specific person based on information about the face, hairstyle, and physique of the specific person included in the image. Furthermore, for example, the first trained model 2023 may be trained to output information about human body parts when an image is input. Specifically, in this embodiment, the first trained model 2023 may identify human body parts such as hands or a face based on the skeletal arrangement of the human included in the image. In this case, the first trained model 2023 identifies human body parts included in the image using a method such as open pose, for example.
[0054] In this case, the learning data may be, for example, a human image as input data, and the judgments about the attributes, specific persons, and human body parts included in the input data may be used as correct output data. In this case, the learning data may not include correct output data.
[0055] The second trained model 2024 is a model trained to output information about an object based on an input image when the image is input. The object here refers to an animal or plant, clothing, machinery, a vehicle, a craft, interior, a building, etc. The second trained model 2024 may be, for example, a separate trained model trained with separate learning data depending on the information to be output.
[0056] For example, the second trained model 2024 learns to output information about an object when an image is input. Specifically, in this embodiment, the second trained model 2024 identifies the object based on information such as the size, shape, material, color, and accessories of the object included in the image.
[0057] In this case, the learning data may be, for example, an image of an object as input data, and a judgment about the object represented in the input data may be the correct output data. In this case, the learning data may not include the correct output data.
[0058] The third trained model 2025 is a model that learns to output information about a movement based on an input image when the image is input. The third trained model 2025 may be, for example, different trained models that are trained with different learning data depending on the information to be output.
[0059] For example, the third trained model 2025 is trained to output information related to shopping when an image is input. Specifically, in this embodiment, the third trained model 2025 is trained to output that a shopping-related activity has been performed when an image containing the following content is input: A customer taking out money or a card near a cash register in a retail store, over multiple frames A customer picking up an item in a retail store in a multi-frame action
[0060] In this case, the learning data may be, for example, an image of the inside of a retail store as input data, and the judgment regarding the behavior represented by this input data may be the correct output data. In this case, the learning data may not include the correct output data.
[0061] The control unit 203 is realized by the processor 29 reading a program stored in the storage unit 202 and executing instructions included in the program. The control unit 203 operates in accordance with the program to perform functions shown as a reception control module 2031, a transmission control module 2032, an image analysis module 2034, an image editing module 2035, and a presentation module 2036.
[0062] The reception control module 2031 controls the process by which the server 20 receives signals from external devices in accordance with a communication protocol.
[0063] The transmission control module 2032 controls the process in which the server 20 transmits signals to external devices in accordance with a communication protocol.
[0064] The image analysis module 2034 analyzes images captured by the imaging device 31 and outputs information about objects at the scene. The image analysis module 2034 inputs the images to the first trained model 2023, the second trained model 2024, and the third trained model 2025. More specifically, the image analysis module 2034 analyzes the image using, for example, the first trained model 2023, and after the analysis by the first trained model 2023 is completed, analyzes the image using the second trained model 2024. After the analysis by the second trained model 2024 is completed, the image analysis module 2034 analyzes the image using the third trained model 2025. The first trained model 2023, the second trained model 2024, and the third trained model 2025 output information about objects corresponding to each model according to the input image. The image analysis module 2034 outputs information output from the first trained model 2023, the second trained model 2024, and the third trained model 2025 to the presentation module 2036. The image analysis module 2034 does not necessarily have to use the first trained model 2023, the second trained model 2024, and the third trained model 2025, and may analyze the image using a trained model specified by the user. The image analysis module 2034 may input an image with a limited analysis range into the trained model. The analysis range may be specified, for example, by a user. Furthermore, time may be added to the conditions for analyzing an image. For example, the image analysis module 2034 analyzes an image only if an object continues to be detected in the analysis range for a certain period of time or more.
[0065] The image editing module 2035 receives editing operations from the user and edits the video analyzed by the image analysis module 2034. Specifically, for example, the image editing module 2035 edits the video from the perspective of time, objects, range on the image, etc. The image editing module 2035 may edit the video in accordance with point-by-point operations by the user, or may automatically edit the video based on rules set by the user.
[0066] The presentation module 2036 presents information about objects in the scene to the user via the terminal device 10. The presentation module 2036 may present information to the user based on an instruction from the user.
[0067] <2 Data Structure> 4 and 5 are diagrams showing the data structure of a database stored in the server 20. Note that FIGS. 4 and 5 are merely examples, and do not exclude data not shown.
[0068] Fig. 4 is a diagram showing the data structure of the imaging device DB 2021. As shown in Fig. 4, each record of the imaging device DB 2021 includes, for example, an item "imaging device ID," an item "coordinates," an item "address," an item "model," an item "installation date," an item "removal date," an item "operation status," an item "latest analysis," an item "current analysis status," an item "analysis frequency," and an item "captured video."
[0069] The item “camera device ID” indicates identification information for identifying the camera device 31 .
[0070] The item "coordinates" indicates the installation location of the camera device 31. Specifically, the item "coordinates" indicates the latitude and longitude of the installation location of the camera device 31. Note that the installation location of the camera device 31 may be determined by indicators other than coordinates.
[0071] The item "Address" indicates the address of the location where the image capturing device 31 is installed.
[0072] The item “model” indicates the model of the imaging device 31 .
[0073] The item "Installation date" indicates the date on which the imaging device 31 was installed.
[0074] The item "Removal Date" indicates the date on which the imaging device 31 was removed.
[0075] The item "operation status" indicates the operation status of the imaging device 31. Specifically, the item "operation status" is expressed as "in operation," "stopped," "out of order," or "removed" according to the status of the imaging device 31.
[0076] The item "Latest Analysis" indicates the time point of the most recent image that has been analyzed by the server 20. Specifically, the server 20 analyzes images received from the imaging device 31 at a predetermined timing, and the item "Latest Analysis" indicates the time point of the most recent image that has been analyzed. In other words, images up to the time point indicated by the item "Latest Analysis" in FIG. 4 (for example, 2023 / 10 / 22 23:59:59) have been analyzed.
[0077] The item "Current Analysis Status" indicates the time division of the image currently being analyzed by the server 20. The time division can be set arbitrarily. For example, the division is as follows: Daily (00:00:00~23:59:59) Morning (00:00:00~9:59:59), Lunch (10:00:00~17:59:59), Dinner (18:00:00~23:59:59)
[0078] The item "analysis frequency" indicates how often the server 20 analyzes images.
[0079] The item "video" indicates a video captured by the camera device 31. Specifically, the video is video data captured by the camera device 31. The video may also store reference information to a video data file located elsewhere.
[0080] Fig. 5 is a diagram showing the data structure of the detection log DB 2022. As shown in Fig. 5, each record in the detection log DB 2022 includes, for example, an item "camera ID," an item "detection model," an item "detection category," an item "detection frame," and an item "image area." When an object is detected by the trained model, a new record is created in the detection log DB 2022, and each piece of information is stored in the corresponding item.
[0081] The item “camera device ID” indicates identification information for identifying the camera device 31 .
[0082] The item "detection model" indicates the trained model that detected the object. Specifically, the first trained model 2023, the second trained model 2024, the third trained model 2025, etc. stored in the storage unit 202 are used as the detection model.
[0083] The item "Detection Category" indicates the category of the detected object. Specifically, for example, the category is a classification of the object detected by the trained model. For example, the object detected by the first trained model 2023 is classified as an "adult male" among humans.
[0084] The item "detection frame" indicates the frame of the image in which the object was detected.
[0085] The item "image area" indicates the area of the object on the image. Specifically, for example, when the image is divided into meshes, the item "image area" indicates the mesh in which the object appears on the image.
[0086] <3 operations> The operation of the server 20 for analyzing images will be described. Although a retail store will be used as an example below, the application of this embodiment is not limited to retail stores. For example, this embodiment may be used to check rule enforcement in a hospital or for safety management in a factory.
[0087] The camera device 31 is attached to the ceiling, for example, from which the equipment in the store can be viewed. In this embodiment, the equipment is, for example, items and locations used for purchasing or selling. In this embodiment, the equipment includes, for example, shelves, entrances and exits, cash registers, and shopping carts. The camera device 31 captures images of the store at a frame rate that allows a series of movements to be recognized as a video. A frame rate that allows a series of movements to be recognized as a video is, for example, about 30 fps. The camera device 31 transmits image data to the server 20.
[0088] The frequency of transmission is a fixed cycle based on the data volume of the unsent video. For example, the image capture device 31 transmits the unsent video to the server 20 every time the data volume of the unsent video reaches a predetermined volume. The frequency of transmission may also be a fixed cycle based on the time of the unsent video. For example, the image capture device 31 transmits the unsent video to the server 20 every time a predetermined time has elapsed since the most recent transmission. Alternatively, the image capture device 31 may immediately transmit the captured video to the server 20.
[0089] FIG. 6 is a flowchart showing the operation when information detected by image analysis is recorded in a log.
[0090] In step S61, the control unit 203 analyzes an image captured by the image capturing device 31 using the first trained model 2023 recorded in the storage unit 202. Specifically, the image analysis module 2034 inputs, for example, image data transmitted from the image capturing device 31 into the first trained model 2023. The frequency at which the image analysis module 2034 analyzes an image is periodic. The image analysis module 2034 may analyze an image, for example, weekly, every few days, daily, or every few hours. The image analysis may be performed, for example, at 9:00 p.m. every Sunday. Alternatively, the image analysis module 2034 may analyze an image immediately after the server 20 receives the image. In other words, the frequency of image analysis by the image analysis module 2034 can be set arbitrarily.
[0091] The first trained model 2023 outputs information related to the detected object.
[0092] In step S62, the control unit 203 determines whether or not the analyzed image contains an object that corresponds to the first trained model 2023. If a corresponding object is present, the control unit 203 shifts the process to step S63. On the other hand, if a corresponding object is not present, the control unit 203 shifts the process to step S64.
[0093] In step S63, the control unit 203 records information related to the object detected in step S62 in the detection log DB 2022. The recorded information includes, for example, the image capture device ID, the detection model, the detection category, the detection frame, and the image area.
[0094] In step S64, if any trained models that have not been used for detection remain, the control unit 203 repeats the processes of steps S61, S62, and S63 using the unused trained models. If any trained models that have not been used for detection remain, the control unit 203 ends the process. When the control unit 203 ends the process, the control unit 203 may notify the display 141 of the terminal device 10 that the control unit 203 has ended the process.
[0095] The user operates the terminal device 10 and accesses the server 20 after user authentication. The server 20 transmits information about the editing screen to the terminal device 10. The editing screen includes information in the detection log DB 2022 that has been updated since the previous access. The user edits the video on the editing screen.
[0096] FIG. 7 is a schematic diagram illustrating an editing screen displayed on the display 141 of the terminal device 10. The display area 71 is an area for displaying video in streaming format. Marks 721, 722, and 723 are borders surrounding objects detected in the video. The style of the borders may differ for each trained model or each detection category. The style of the borders may include, for example, color, thickness, shape, etc., or a combination thereof. Even if an object of only one detection model or only one detection category is detected in the video, a mark may be attached to the object. The number 73 indicates the time point in the video that is being displayed. The icon group 74 is an icon for operating the video. The icon group 74 is not limited to the three icons shown in FIG. 7.
[0097] The timeline 75 is a time axis for displaying the time associated with the detection marker 78 of each box 77 on a timeline. The timeline 75 is divided into sections at regular time points, for example. The playback marker 76 is a marker on the timeline 75 that indicates the time point of the video currently being played. The box 77 is a box that indicates the category of the detected object. The box 77 may also indicate information that can identify the trained model of the detected object. The detection marker 78 is a marker on the timeline 75 that indicates the time at which the object was detected. According to the example of FIG. 7, the dog was detected between 8:00 and 9:00.
[0098] Icon 79 is an icon for editing a video. For example, icon 79 is an icon for dividing a video into certain time segments and then cutting out the video. For example, a user may determine the time of the video to be cut out by specifying a detection marker 78. That is, the image editing module 2035 cuts out the video of the time period corresponding to the detection marker 78. The image editing module 2035 is not limited to matching the time period of the video to be cut out with the detection marker 78, and may add a predetermined time width to the detection marker 78. For example, the image editing module 2035 may add a predetermined time width before the start point of the detection marker 78, after the end point of the detection marker 78, or both. Furthermore, the user may determine the time of the video to be cut out by, for example, specifying a time range on the timeline 75. Specifically, for example, the image editing module 2035 receives a tap or a drag at two points on the timeline 75 from the user, and cuts out the video of the received time range.
[0099] After editing the video, the image editing module 2035 saves the edited video file in, for example, the memory 25 or the storage 26.
[0100] FIG. 8 is a schematic diagram showing a portion of the editing screen displayed on the display 141 of the terminal device 10. As shown in FIG. 8, the user may determine the time of the video to be cut out by, for example, specifying two or more detection markers. Specifically, as shown in editing case 81, the user may determine the time of the video to be cut out by specifying two detection markers that are separated from each other on the timeline. In other words, the image editing module 2035 may cut out the video in the time periods corresponding to the detection markers specified by the user and combine the videos. The two specified detection markers may belong to the same category or different categories.
[0101] Specifically, as shown in editing case 82 and editing case 83, the user may determine the time of the video to be cut out by specifying two overlapping detection markers on the timeline. In editing case 82, the image editing module 2035 extracts the portion where the two detection markers specified by the user overlap. In editing case 83, the image editing module 2035 extracts the portion corresponding to at least one detection marker.
[0102] The icon 79 in Fig. 7 is not limited to the icon for cutting out shown in Fig. 7. For example, the icon 79 may be an icon for editing a moving image. Specifically, the icon 79 may be, for example, an icon for adjusting the brightness of a moving image or an icon for applying a predetermined effect to the moving image. In other words, the image editing module 2035 may accept an operation on the icon 79 from the user and edit the moving image.
[0103] The user may arbitrarily set rules for editing the video. Specifically, the user may set rules for automatic editing based on an object as a condition. For example, the user may set rules for automatically determining the time of the video to be cut out by specifying a child detection marker 78 and a toy detection marker 78. As a result, when a predetermined object is detected, a video including the object is automatically generated. Specifically, the user may set rules for automatic editing based on a time as a condition. For example, the user may set a rule for automatically increasing the brightness of a video shot at night. As a result, a video is automatically generated during a predetermined time period. In other words, the image editing module 2035 may automatically edit the video according to any video editing rules set by the user.
[0104] Icon 710 in Fig. 7 is an icon for downloading the video edited by icon 79. In response to a user's operation on icon 710, the transmission control module 2032 transmits the edited video file from the memory 25, the storage 26, or the like to the terminal device 10. The terminal device 10 saves the downloaded video file.
[0105] As described above, in the above embodiment, the control unit 203 acquires a captured video. The control unit 203 analyzes the video using a plurality of pre-stored analytical models, each of which has been learned for a target object. The control unit 203 presents to the user a detection marker 78 indicating the timing at which the target object was detected by the analytical model, in association with the timeline 75 associated with the video. The control unit 203 generates an edited video from the captured video in response to a user instruction based on the detection marker 78. This allows the user to intuitively and easily grasp the situation at the scene where the video was captured via a user interface in which the appearance times of objects for each model are visually organized on the timeline 75. This also allows the user to intuitively and easily edit the video via a user interface for efficiently processing objects.
[0106] Therefore, the system according to this embodiment improves convenience when editing videos based on images captured by a surveillance camera or the like.
[0107] In the above embodiment, the control unit 203 receives a specification for the detection marker 78 in the generating step, and generates an edited video for the time period corresponding to the detection marker 78 or a time period with a predetermined range around the time period. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object disappears.
[0108] In the above embodiment, the control unit 203 receives designations for multiple detection markers 78 in the generating step, and generates an edited video of a combined time period by joining together time periods corresponding to the detection markers 78, or a combined time period by joining together time periods with a predetermined width therebetween. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object disappears.
[0109] In the above embodiment, if there is a time period in which the detection markers 78 detected by different analysis models overlap in the generating step, the control unit 203 generates an edited video for the overlapping time period or for a time period that has a predetermined width within the time period. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object disappears.
[0110] In the above embodiment, if there is a time period in which detection markers 78 based on different analysis models overlap in the generating step, the control unit 203 generates an edited video for the time period corresponding to at least one of the detection markers 78, or for a time period with a predetermined width around that time period. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object disappears.
[0111] In the above embodiment, the control unit 203 receives a designation of a time period on the timeline 75 in the generating step, and generates an edited video for the designated time period or a combined time period by joining together the designated time periods. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to edit the video without being limited by the detection markers 78, while still referring to the detection markers 78.
[0112] In the above embodiment, in the generating step, the control unit 203 receives instructions from the user specifying rules for generating an edited moving image based on the detected markers 78. This allows the user to edit the moving image more intuitively and simply via a user interface for efficiently processing objects. This also saves the user the trouble of inputting frequently occurring editing patterns.
[0113] In the above embodiment, the control unit 203 analyzes the video at an arbitrary frequency using the analysis model in the analysis step. As a result, the frequency of image analysis by the image analysis module 2034 can be set arbitrarily, so the device according to this embodiment does not necessarily require high performance.
[0114] In the above embodiment, in the presenting step, when objects detected by different analysis models are present in the same frame, the control unit 203 simultaneously presents marks 721 to 723 representing the objects to the user. This allows the user to intuitively and easily grasp the situation of the scene photographed by the marks 721 to 723 attached to each object, even when multiple objects are detected in the same frame.
[0115] In the above embodiment, the control unit 203 presents the marks 721 to 723 in different forms for each analysis model in the presenting step, which allows the user to intuitively and easily grasp the situation of the captured scene from the marks 721 to 723 that differ in form for each analysis model.
[0116] <4 Other embodiments> In the above embodiment, it has been described that the image capturing device 31 transmits video to the server 20. However, the server 20 may be realized by a general information processing device such as a PC. In this case, for example, a storage medium storing video in the image capturing device 31 is removed and connected to an input port of the information processing device, thereby inputting the information stored in the storage medium to the information processing device. The control unit of the information processing device performs processing similar to that of the image analysis module 2034 of the server 20, for example, analyzing images captured by the image capturing device 31 and outputting information about objects at the site. In addition, the control unit of the information processing device performs processing similar to that of the image editing module 2035 of the server 20, for example, accepting editing instructions from a user and generating an edited video based on the accepted editing instructions.
[0117] In the above embodiment, it has been described that the edited video is downloaded from the memory 25, the storage 26, or the like to the terminal device 10 and stored therein. However, the user may view the edited video on the server 20 without downloading it.
[0118] In the above embodiment, it has been described that the image analysis module 2034 may limit the analysis range of an image. The analysis range for one image may be two or more. For example, the ranges analyzed for a video taken of a retail store may be two locations: the entrance and the exit. Furthermore, if objects are detected simultaneously in two or more analysis ranges, or if objects are detected within a certain time interval, the video for the times when the objects were detected may be extracted.
[0119] In the above embodiment, it has been described that a video of a time period equal to the detected marker plus a predetermined time width may be cut out. However, the video to be created is not limited to this. The image editing module 2035 may cut out a video of a time period consisting of only the predetermined time width.
[0120] In the above embodiment, the editing screen displays one video captured by one camera device 31 and one timeline corresponding to that video. However, what is displayed on the editing screen is not limited to these. In other words, when generating an edited video, the videos referenced are not limited to those captured by one camera device 31. For example, the image analysis module 2034 may analyze videos captured by two or more camera devices 31. Furthermore, the image editing module 2035 may generate an edited video based on videos captured by two or more camera devices 31. Specifically, for example, the image analysis module 2034 analyzes multiple videos captured by multiple camera devices 31 installed to capture the same location from different angles. The image editing module 2035 generates an edited video based on the analysis results.
[0121] FIG. 9 is a schematic diagram illustrating another example of an editing screen displayed on the display 141 of the terminal device 10. Note that while FIG. 9 illustrates a case where one object is detected, multiple objects may be detected. In other words, multiple trained models may be used in the analysis. The editing screen in FIG. 9 includes, for example, multiple videos and one timeline collectively corresponding to the videos.
[0122] In Figure 9, detection marker 78 indicates that an object has been detected in the video. As indicated by detection marker 78 in Figure 9, presentation module 2036 may display a detection marker for each video from each camera under the same object. By manipulating multiple detection markers under the same object, a user can extract videos based on, for example, the time at which the same object was detected in multiple videos.
[0123] As shown by the display area 71 in FIG. 9, multiple videos are displayed in divided display areas within the editing screen.
[0124] The image editing module 2035 may, for example, provide the user with options for the videos to be edited. Specifically, for example, the image editing module 2035 may accept, from the user, a selection of videos to be included in the edited video. For example, when the user selects one or more detection markers under the same object, the image editing module 2035 sets the videos corresponding to the selected detection markers as the editing target. The user may also select a video in which no detection marker appears under the same object.
[0125] If the user selects a single video, the image editing module 2035 edits the selected video as if the editing screen were displaying one video and one timeline corresponding to that video.
[0126] If the user selects multiple videos, the image editing module 2035 edits the selected videos, specifying how the multiple videos should be played, i.e., whether the selected videos should be played in sequence or in parallel.
[0127] When multiple videos are played in sequence, the order of playback is determined automatically or by user operation. The automatically determined order of playback may be determined, for example, based on time, the proportion of the screen, etc. Specifically, the image editing module 2035 determines the order of playback taking into consideration, for example, the length of time indicated by the detection marker in the video, the average proportion of the area of the object that occupies the screen throughout the video, etc. Furthermore, when the user determines the order of playback, the image editing module 2035 may suggest the automatically determined order to the user as an aid in the decision.
[0128] When multiple videos are played in parallel, the multiple videos are played, for example, in a picture-in-picture format in which the secondary video is placed on top of the primary video in the video playback area. The primary-subordinate relationship between the videos is determined automatically or by user operation. The automatically determined primary-subordinate relationship may be determined based on, for example, time, the proportion of the screen occupied, etc. Specifically, the image editing module 2035 determines the primary-subordinate relationship by taking into consideration, for example, the length of time indicated by the detection marker in the video, the average proportion of the area of the object occupied on the screen throughout the video, etc. Furthermore, when the user determines the primary-subordinate relationship, the image editing module 2035 may suggest the automatically determined primary-subordinate relationship to the user as an aid in the determination.
[0129] Furthermore, when multiple videos are played in parallel, the multiple videos may be played in, for example, divided video playback areas. That is, the image editing module 2035 divides the video playback area so that multiple videos are played in the video playback area. The area division is determined automatically or in response to a user operation. The automatically determined division may be determined based on, for example, time, the proportion of the screen occupied, etc. Specifically, the image editing module 2035 determines the division taking into account, for example, the length of time indicated by the detection marker in the video, the average proportion of the area of the object occupied on the screen throughout the video, etc. Furthermore, when the user determines the division, the image editing module 2035 may suggest the automatically determined division to the user as an aid in the decision. Even if there are differences between the multiple videos in terms of time, the proportion of the screen occupied, etc., the image editing module 2035 may divide the video playback area evenly in response to a user operation.
[0130] In the above embodiment, a case has been described in which a detection marker relating to a detected object appears on the timeline, but information other than the detected object may also appear on the timeline. For example, the presentation module 2036 may display output information of a predetermined IoT device on a timeline. Specifically, the reception control module 2031 receives information measured by an IoT device attached to on-site equipment at a predetermined interval. The presentation module 2036 presents a detection marker based on the information transmitted from the IoT device on the timeline of the editing screen. For example, if the acquired information satisfies predetermined requirements, the presentation module 2036 presents the detection marker on the timeline of the editing screen. For example, when the reception control module 2031 receives information indicating a temperature above a certain level from a temperature sensor attached to equipment displayed on an image, the presentation module 2036 presents a detection marker on a timeline indicating the time during which that temperature was reached. The presentation module 2036 may present the detection markers so that the detection markers indicate temperature ranges in different ways for each temperature range. For example, an orange detection marker may indicate a temperature between 30 degrees Celsius and 40 degrees Celsius. Furthermore, for example, when the reception control module 2031 receives information about acceleration equal to or greater than a certain level from an acceleration sensor attached to equipment shown on an image, the presentation module 2036 presents a detection marker on a timeline indicating the time during which that acceleration occurred. The presentation module 2036 may present the detection marker so that the detection marker indicates the acceleration range in a different manner for each acceleration range. For example, a blue detection marker may indicate an acceleration of 1 G or greater but less than 2 G.
[0131] <5. Basic computer hardware configuration> 10 is a block diagram showing the basic hardware configuration of a computer 100. The computer 100 includes at least a processor 101, a main memory device 102, an auxiliary memory device 103, and a communication IF (interface) 109. These are electrically connected to each other by a bus.
[0132] The processor 101 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0133] The main storage device 102 is used to temporarily store programs, data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0134] The auxiliary storage device 103 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.
[0135] The communication IF 109 is an interface for inputting and outputting signals for communicating with other computers via a network using a wired or wireless communication standard.
[0136] The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In the case of a wired connection, the network also includes a direct connection using a USB (Universal Serial Bus) cable, etc.
[0137] It should be noted that the computer 100 can be virtually realized by distributing all or part of each hardware configuration across multiple computers 100 and interconnecting them via a network. In this way, the computer 100 is a concept that includes not only a computer 100 housed in a single housing or case, but also a virtualized computer system.
[0138] <Basic functional configuration of computer 100> A description will be given of the functional configuration of a computer realized by the basic hardware configuration of the computer 100 shown in Fig. 10. The computer includes at least the functional units of a control unit, a storage unit, and a communication unit.
[0139] The functional units of the computer 100 can also be realized by distributing all or part of the functional units among multiple computers 100 interconnected via a network. The computer 100 is a concept that includes not only a single computer 100 but also a virtualized computer system.
[0140] The control unit is realized by the processor 101 reading various programs stored in the auxiliary storage device 103, expanding them in the main storage device 102, and executing processing in accordance with the programs. The control unit can realize functional units that perform various types of information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0141] The storage unit is realized by a main storage device 102 and an auxiliary storage device 103. The storage unit stores data, various programs, and various databases. Furthermore, the processor 101 can allocate a storage area corresponding to the storage unit in the main storage device 102 or the auxiliary storage device 103 in accordance with the programs. Furthermore, the control unit can cause the processor 101 to execute processes for adding, updating, and deleting data stored in the storage unit in accordance with the various programs.
[0142] A database refers to a relational database, which manages data sets called tables, which are structured by rows and columns, by relating them to each other. In a database, a table is called a table, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables can be set and associated. Typically, each table has a column set as a key for uniquely identifying a record, but setting a key to a column is not essential. The control unit can cause the processor 101 to add, delete, or update records in a specific table stored in the storage unit according to various programs.
[0143] The communication unit is realized by the communication IF 109. The communication unit realizes a function of communicating with other computers 100 via a network. The communication unit can receive information transmitted from other computers 100 and input the information to the control unit. The control unit can cause the processor 101 to execute information processing on the received information in accordance with various programs. Furthermore, the communication unit can transmit information output from the control unit to other computers 100.
[0144] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to perform the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes programs stored in memory. In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions. If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0145] Although several embodiments of the present disclosure have been described above, these embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and modifications are intended to be included in the scope of the inventions and their equivalents as defined in the claims, as well as in the scope and spirit of the inventions.
[0146] <Additional Notes> The matters described in the above embodiments will be supplemented below. (Appendix 1) A program for operating a computer having a processor 29 and a memory 25, the program causing the processor to execute the steps of acquiring a filmed video, analyzing the acquired video using a plurality of pre-stored analytical models, each of which has been trained to detect an object, presenting to the user a detection marker 78 indicating the timing at which the object was detected by the analytical model, in association with a timeline 75 associated with the video, and generating an edited video from the filmed video in response to instructions from the user based on the detection marker. (Appendix 2) The program described in (Appendix 1), wherein in the generating step, a specification for a detection marker is accepted, and an edited video is generated for a time period corresponding to the detection marker or a time period with a predetermined range around that time period. (Appendix 3) A program described in (Appendix 1) or (Appendix 2), in which, in the generating step, specifications for multiple detection markers are accepted, and an edited video of a combined time period is generated by connecting time periods corresponding to the detection markers, or by connecting time periods with a predetermined width within the time period. (Appendix 4) A program described in any one of (Appendix 1) to (Appendix 3), wherein in the generating step, if there is a time period in which detected markers by different analysis models overlap, an edited video is generated for the overlapping time period or for a time period with a predetermined width within that time period. (Appendix 5) A program described in any one of (Appendix 1) to (Appendix 4), wherein in the generating step, if there is a time period in which detection markers from different analysis models overlap, an edited video is generated for a time period corresponding to at least one of the detection markers, or for a time period with a predetermined range around that time period. (Appendix 6) A program described in any one of (Appendix 1) to (Appendix 5), wherein in the generating step, a specification of a time period on a timeline is accepted and an edited video of the specified time period or a combined time period created by connecting these time periods is generated. (Appendix 7) The program according to any one of (Supplementary Note 1) to (Supplementary Note 6), wherein in the generating step, a rule for generating an edited video based on the detected markers is given as an instruction from a user. (Appendix 8) The program according to any one of (Appendix 1) to (Appendix 7), wherein in the analyzing step, the video is analyzed at an arbitrary frequency using an analysis model. (Appendix 9) A program described in any one of (Appendix 1) to (Appendix 8), wherein in the presentation step, if objects detected by different analysis models are in the same frame, marks representing the objects are simultaneously presented to the user. (Appendix 10) The program according to (Appendix 9), wherein in the presenting step, the mark has a different form for each analysis model. (Appendix 11) A method executed by a computer having a processor and a memory, wherein the processor executes all of the steps performed in any of the inventions according to (Appendix 1) to (Appendix 10). (Appendix 12) An information processing device comprising a control unit 203 and a storage unit 202, wherein the control unit executes all of the steps executed in the invention according to any one of (Appendix 1) to (Appendix 10). (Appendix 13) A system comprising means for executing all steps performed in any of the inventions according to (Appendix 1) to (Appendix 10). [Explanation of symbols]
[0147] 1. System 10...Terminal device 12...Communication IF 120…Communications Department 13...Input device 131...Touch-sensitive devices 14...Output device 141...Display 15...Memory 150...Location information sensor 16…Storage 160...camera 17...Audio processing unit 171...Mike 172...Speaker 180...Storage section 181...User information 19...Processor 190...Control unit 191...Operation reception section 192...Transmitter / receiver 193...Presentation control unit 20...Server 201…Communications Department 202...Storage section 2021…Photography Equipment DB 2022…Detection log database 2023…First trained model 2024…Second trained model 2025…Third trained model 203...Control unit 2031...Receiver control module 2032...Transmission control module 2034...Image analysis module 2035...Image editing module 2036… Presentation module 22...Communication IF 23...Input / output interface 25…Memory 26…Storage 29...Processor 31...Photographic equipment 80…Network 100...Server (computer) 101...Processor 102...Memory (main storage device) 103...Storage (auxiliary storage device) 109…Communication IF
Claims
1. A program for operating a computer comprising a processor and memory, wherein the program is configured to operate the processor, Steps to acquire the recorded video, The steps include: analyzing the acquired video using multiple pre-stored analysis models, each of which has been trained to detect an object; The steps include presenting to the user, in chronological order, information indicating the timing at which the object was detected by the analysis model in the aforementioned video, The steps include generating an edited video from the recorded video in response to instructions from the user based on the information provided, and Make it run, In the generation step, if there is a time period in which the timings indicated by the information from different analysis models overlap, the program generates the edited video for the overlapping time period, or for a time period with a predetermined width within that time period.
2. A program for operating a computer comprising a processor and memory, wherein the program provides the processor, Steps to acquire the recorded video, The steps include: analyzing the acquired video using multiple pre-stored analysis models, each of which has been trained to detect an object; The steps include presenting to the user, in chronological order, information indicating the timing at which the object was detected by the analysis model in the aforementioned video, The steps include generating an edited video from the recorded video in response to instructions from the user based on the information provided, and Make it run, In the generation step, if there is a time period in which the timings indicated by the information from different analysis models overlap, the program generates the edited video for a time period corresponding to the timing indicated by at least one of the pieces of information, or for a time period with a predetermined width within that time period.
3. A program for operating a computer comprising a processor and memory, wherein the program provides to the processor, Steps to acquire the recorded video, The steps include: analyzing the acquired video using multiple pre-stored analysis models, each of which has been trained to detect an object; The steps include presenting to the user, in chronological order, information indicating the timing at which the object was detected by the analysis model in the aforementioned video, The steps include generating an edited video from the recorded video in response to instructions from the user based on the information provided, and Make it run, A program that, in the generation step, is given by the user, as instructions, rules for generating the edited video based on the information provided.
4. A program for operating a computer comprising a processor and memory, wherein the program provides the processor, Steps to acquire the recorded video, The steps include: analyzing the acquired video using multiple pre-stored analysis models, each of which has been trained to detect an object; The steps include presenting to the user, in chronological order, information indicating the timing at which the object was detected by the analysis model in the aforementioned video, The steps include generating an edited video from the recorded video in response to instructions from the user based on the information provided, and Make it run, A program that, in the aforementioned steps, simultaneously presents to the user a marker representing an object if an object detected by a different analysis model is present in the same frame.
5. The program according to claim 4, wherein in the steps described above, the mark has a different form for each analysis model.
6. A method to be performed on a computer comprising a processor and memory, wherein the processor performs all steps performed in any of the inventions according to claims 1 to 5.
7. An information processing apparatus comprising a control unit and a storage unit, wherein the control unit performs all steps performed in the invention according to any one of claims 1 to 5.
8. A system comprising means for performing all steps performed in the invention according to any one of claims 1 to 5.