Video aggregation computer, video aggregation method, and program
The video aggregation computer addresses data transfer inefficiencies by synthesizing and compressing images from multiple observation devices into a single image for remote monitoring, enhancing efficiency and adaptability in varying communication conditions.
Patent Information
- Application Number
- PCT/JP2025/022809
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2025-06-25
- Publication Date
- 2026-01-02
AI Technical Summary
Existing systems for remotely operating work machines at dangerous sites, such as construction sites, face challenges in managing the large volume of data transfer from multiple imaging devices, leading to inefficiencies and potential communication issues.
A video aggregation computer that acquires images from multiple observation devices, sets specific video portions for display, synthesizes these portions into a single image, and transmits this image to a terminal, reducing data transfer volume and optimizing communication.
This approach reduces data transfer requirements, allowing efficient monitoring of multiple sites with minimal infrastructure, even in environments with limited bandwidth, by synthesizing and compressing video data for seamless display on remote terminals.
Smart Images

Figure JP2025022809_02012026_PF_FP_ABST
Abstract
Description
Video aggregation computer, video aggregation method, and program
[0001] The present invention relates to a video aggregation computer, a video aggregation method, and a program.
[0002] Conventionally, at particularly dangerous sites such as construction sites, work machines and the like are sometimes remotely operated. At such sites, imaging devices such as fixed cameras installed at multiple locations, cameras mounted on drones, and mobile cameras installed on portable devices, etc., are used to provide work support to workers, managers, and other related parties in various remote locations with information about the site situation (e.g., Patent Documents 1 and 2).
[0003] International Publication No. 2021 / 070214 Pamphlet Patent No. 7029586 Specification
[0004] An object of the present invention is to provide a video aggregation computer, a video aggregation method, and a program that can reduce the amount of data transfer.
[0005] The provided video aggregation computer is a video aggregation computer for displaying images from multiple distributed observation devices on a terminal, and includes an acquisition unit that acquires multiple images taken by the multiple observation devices, a setting unit that sets the image portion to be displayed on the terminal for each of two or more images from the multiple acquired images, a synthesis unit that synthesizes the two or more set image portions into a single image, and an output unit that transmits the synthesized single image to the terminal.
[0006] FIG. 1 is a diagram for explaining an overview of a system including a video aggregation computer 1, an observation device 2, and a terminal 3 according to an embodiment of the present invention. FIG. 2 is a configuration diagram of a system including a video aggregation computer 1, an observation device 2, and a terminal 3 according to this embodiment. FIG. 3 is a flowchart of a video aggregation process executed by the video aggregation computer 1 according to this embodiment. FIG. 4 is a diagram for explaining an example of a video that the video aggregation computer 1 according to this embodiment displays (outputs) on the terminal 3.
[0007] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, a detailed description of the present invention will be given with reference to the accompanying drawings. In the drawings, the same elements are designated by the same numbers or symbols throughout the description of the embodiments.
[0008] [Basic Concept / Basic Configuration] Fig. 1 is a diagram for explaining an overview of a video aggregation system that is a system including a video aggregation computer 1 according to one embodiment of the present invention, a plurality of observation devices 2, and a terminal (information terminal) 3. Hereinafter, the video aggregation computer 1 may be referred to simply as computer 1.
[0009] As shown in FIG. 1 , the video consolidation computer 1 may function as, for example, a media server. The computer 1 is connected to each of the observation device 2 and the terminal 3 for data communication via the Internet, a network such as a mobile phone network, a network constructed by software virtualization, or the like. The number of each of the observation device 2 and the terminal 3 may be one or more. Two or more observation devices 2 may be connected to the computer 1 simultaneously. The observation device 2 and the terminal 3 may be connected to the computer 1 simultaneously. The system shown in FIG. 1 includes multiple observation devices 2. The multiple observation devices 2 may be installed at a single site or may be installed separately at multiple sites. The observation device 2 may be, for example, an imaging device such as a video camera that observes the remote operation status of a remotely controlled work machine such as a robot arm, shovel, crane, bulldozer, or transporter. The multiple observation devices 2 may include an observation device 2 mounted on a mobile object. The mobile object may be a mobile work machine such as a shovel or crane, a vehicle such as an automobile, or an aerial vehicle such as a drone. The terminal 3 acquires and displays the images captured by the observation device 2 via the image aggregation computer 1.
[0010] The video aggregation computer 1 may be an on-premise computer or computing system such as an on-premise server or on-premise computing system, or may be a cloud computer or computing system such as a cloud server or cloud computing system. In this embodiment, the video aggregation computer 1 is a cloud computing system. As will be described later, the computer 1 may be a personal computer, a computer installed in a mobile terminal, or a computer installed in a wearable terminal.
[0011] The observation device 2 may be a device such as a camera, a temperature sensor, or a metal detector. The video captured by the observation device 2 may be video data, image data, numerical data, or other data. The video aggregation computer 1 may acquire the data from the observation device 2 via the network, or may acquire the data via a system (not shown) that controls remote operations other than the observation device 2.
[0012] The terminal 3 is a terminal capable of transmitting and receiving data between the terminal 3 and the video aggregation computer 1. The terminal 3 may be, for example, an electronic device such as a laptop computer, a desktop computer, a smartphone, or a tablet terminal.
[0013] As shown in Fig. 1, the computer 1 has at least one video channel 4. The video channel 4 includes a plurality of ports 5. Each of the plurality of ports 5 may be assigned a unique port number (identifier). The video channel 4 may be, for example, a data channel in a peer-to-peer communication connection method. An example of communication based on a peer-to-peer communication connection method is WebRTC (Web Real-Time Communication).
[0014] The video aggregation computer 1 acquires multiple videos captured by multiple observation devices 2. Specifically, the video aggregation computer 1 may acquire multiple videos from multiple observation devices 2 observing the site X via a video channel 4. The video channel 4 may be composed of multiple ports 5 set according to the multiple observation devices 2. The video aggregation computer 1 may acquire video from each observation device 2 via the port 5 corresponding to that observation device 2. The multiple observation devices 2 may be, for example, a portion of the multiple observation devices 2 present at the site X shown in FIG. 1 , or may be all of the multiple observation devices 2.
[0015] The video aggregation computer 1 sets video portions to be displayed on the terminal 3 for each of two or more videos among the multiple videos acquired from the multiple observation devices 2. The video aggregation computer 1 may set video portions to be displayed on the terminal 3 for each of the multiple videos acquired from the multiple observation devices 2. The video aggregation computer 1 may set video portions required by the terminal 3 for each of the multiple acquired videos. That is, the video portions to be displayed on the terminal 3 may be video portions required by the terminal 3. The video portions required by the terminal 3 may be video portions required by a user, such as the user of the terminal 3. The video portions to be displayed on the terminal 3 may be video portions that the user wants to display on the terminal 3, i.e., video portions that the user desires to be displayed on the terminal 3. Furthermore, the video portions to be displayed on the terminal 3 may be video portions targeted for display on the terminal 3 based on information such as information from the observation devices 2, inputs made by the user to the terminal 3, inputs made by the user to the computer 1, and inference results inferred using a trained model, as will be described later.
[0016] The image portion is at least a part of the image (original image) that the computer 1 acquires from the observation device 2. The image portion is a part or all of the image (original image) that the computer 1 acquires from the observation device 2.
[0017] The video portion may be represented, for example, by a specific range (specific area) in the original video. The specific range (specific area) may be represented, for example, by a specific coordinate range in the original video. The specific range (specific area) may be, for example, a range (area) in the original video that corresponds to a specific work area. The specific range (specific area) may be, for example, a range (area) in the original video that includes a specific object. The specific object may be, for example, an object present at the work site, a person such as a worker present at the work site, or another object. The object present at the work site may be, for example, a work machine present at the work site, an obstacle present at the work site, or equipment present at a manufacturing site.
[0018] The video aggregation computer 1 may set the portion of the video to be displayed on the terminal 3 (e.g., the portion of the video required by the terminal 3 or the user) for each acquired video. The computer 1 may set the portion of the video to be displayed on the terminal 3 (e.g., the portion of the video required by the terminal 3) based on information from the observation device 2, based on input provided to the terminal 3 by a remote user of the terminal 3, or based on an inference result inferred using a trained model constructed by machine learning. The video aggregation computer 1 may set the portion of the video required by the terminal 3, for example, based on information for specifying the range (area) of the video required by the terminal 3. Specific examples of setting the video portion will be described later.
[0019] The video aggregation computer 1 combines the multiple video portions set for the multiple videos to synthesize the multiple video portions into a single video. The video aggregation computer 1 may trim at least one of the multiple acquired videos. Trimming is a process (cutting process) in which the computer 1 retains the video portions that are required by the terminal 3 from the video (original video) acquired by the observation device 2 and removes unnecessary portions from the original video. Specifically, for example, the computer 1 may generate the video portions by trimming the original video based on the settings for the video portions. The computer 1 may combine the generated video portions to synthesize the multiple video portions into a single video. The computer 1 may combine the generated video portions into a single video and compress the combined single video. The computer 1 may also compress each of the generated video portions and combine the compressed video portions to synthesize the multiple video portions into a single video. Specific examples of the settings will be described later.
[0020] The video aggregation computer 1 transmits the single composite video to the terminal 3. Specifically, for example, the video aggregation computer 1 may transmit the single composite video to the terminal 3 and output it as a single screen on the terminal 3. The video aggregation computer 1 may transmit video data for the single composite video to the terminal 3, and the terminal 3 may use the received video data to display the single video synthesized by the video aggregation computer 1 on a single screen of the display device of the terminal 3, as in the display example of the terminal 3 in FIG.
[0021] 2 is a configuration diagram of a video aggregation system according to this embodiment, which is a system including a video aggregation computer 1, an observation device 2, and a terminal 3. The video aggregation computer 1 may be realized, for example, by one terminal device or by multiple terminal devices.
[0022] As described above, the computer 1 may be an on-premise computer or computing system, or a cloud-based computer or computing system. The computer 1 may also be a personal computer such as a desktop computer or a laptop computer. The video aggregation computer 1 may also be a computer installed in a mobile terminal such as a handheld terminal, smartphone, or tablet terminal, or a computer installed in a wearable terminal such as smart glasses, a head-mounted display, or a smart watch. In this case, the computer 1 may be equipped with a camera or other imaging device that captures color video and / or still images.
[0023] The video-intensive computer 1 includes an arithmetic processing unit and a memory. Examples of the arithmetic processing unit include a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). Examples of the memory include a RAM (Random Access Memory) and a ROM (Read Only Memory). The computer 1 includes a control unit. The functions of the control unit are realized by the arithmetic processing unit executing a program stored in the memory. As shown in FIG. 2 , the control unit includes a processing unit, a communication unit, a storage unit, an input unit, and an output unit. The control unit issues execution commands to the processing unit, communication unit, input unit, output unit, storage unit, etc., and the processing unit calculates data and determines the calculation results.
[0024] The video aggregation computer 1 includes a communication unit that is a device that enables the computer 1 to communicate with other devices such as the terminal 3 and the observation device 2. The communication method may be wireless or wired.
[0025] The video-intensive computer 1 has, as an input unit, functions necessary for a user to operate the video-intensive computer 1. The computer 1 includes an input device for realizing input. The computer 1 can be equipped with, for example, an LCD display that realizes a touch panel function, a keyboard, a mouse, a pen tablet, hardware buttons on the device, a microphone for voice recognition, and the like as input devices. The input unit of the computer 1 is not limited to the input methods described above.
[0026] The video-intensive computer 1 has, as an output unit, functions necessary for a user to operate the video-intensive computer 1. The computer 1 includes an output device for realizing output. The computer 1 may, for example, be equipped with a display device or an audio output device as the output device. Examples of the display device include a liquid crystal display, a PC display, a projector, a head-mounted display, etc., and examples of the audio output device include a speaker, etc. The output unit of the computer 1 is not limited to the output methods described above.
[0027] The video aggregation computer 1 includes a data storage (recording medium) such as a hard disk, semiconductor memory, or memory card as a memory unit. The data may be stored in a cloud service, a database, or the like. The memory unit may store, for example, some or all of the video acquired from the observation device 2 as data. In this case, the computer 1 may be configured to be able to search for video data stored in the memory unit.
[0028] The processing unit of the control unit includes a setting unit 10 and a synthesis unit 11. The communication unit of the control unit includes an acquisition unit 20, an output unit 21, and a reception unit 22. The storage unit of the control unit includes a video storage unit 30.
[0029] The observation device 2 may be, for example, a camera, a temperature sensor, a metal detector, or the like. The observation device 2 is equipped with a camera unit that captures video. The observation device 2 may also be, for example, a mobile terminal such as a handheld terminal, smartphone, or tablet terminal, or a wearable terminal such as smart glasses, a head-mounted display, or a smartwatch. The multiple observation devices 2 shown in FIG. 1 may include an observation device 2 attached to a mobile body for observing the remote operation status of the work machine, or may include an observation device 2 attached to a work machine. Two or more observation devices 2 may be attached to one mobile body, and two or more observation devices 2 may be attached to one work machine. In this case, the two or more observation devices 2 are disposed at positions distant from each other on the mobile body or the work machine. The multiple observation devices 2 shown in FIG. 1 may also include one or more observation devices 2 installed on-site for observing the remote operation status of the work machine. The multiple observation devices 2 shown in FIG. 1 may also include one or more observation devices 2 installed on-site for observing the remote operation status of the mobile body. The multiple observation devices 2 shown in Figure 1 include one or more observation devices 2 attached to a work machine, an observation device 2 attached to a mobile body, and one or more observation devices 2 fixed to the work site.
[0030] The observation device 2 includes a processing unit and a memory. Examples of the processing unit include a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). Examples of the memory include a RAM (Random Access Memory) and a ROM (Read Only Memory). The observation device 2 includes a control unit. The functions of the control unit are realized by the processing unit executing a program stored in the memory. As shown in FIG. 2 , the control unit of the observation device 2 includes a processing unit, a communication unit, a storage unit, an imaging unit, an input unit, and an output unit. The control unit issues execution commands to the processing unit, imaging unit, communication unit, input unit, output unit, and storage unit, and the processing unit calculates data and determines the calculation results.
[0031] The observation device 2 includes a communication unit that is a device that enables the observation device 2 to communicate with other devices such as the computer 1 and the terminal 3. The communication method may be wireless or wired.
[0032] The observation device 2 has an input unit that has the functions necessary for a remote user to operate the observation device 2. The observation device 2 includes an input device for realizing input. The observation device 2 can be equipped with, for example, an LCD display that realizes a touch panel function, a keyboard, a mouse, a pen tablet, hardware buttons on the device, a microphone for voice recognition, etc. as input devices. The input unit of the observation device 2 is not limited to the input methods described above.
[0033] The observation device 2 has, as an output unit, functions necessary for a remote user to operate the observation device 2. The observation device 2 includes an output device for realizing output. The observation device 2 may, for example, be equipped with a display device or an audio output device as the output device. Examples of display devices include a liquid crystal display, a PC display, a projector, etc., and examples of audio output devices include a speaker, etc. The output unit of the observation device 2 is not limited to the output methods described above.
[0034] The observation device 2 includes a storage unit such as a hard disk, a semiconductor memory, a memory card, or other data storage (recording medium). The data may be stored in a cloud service, a database, or the like.
[0035] Terminal 3 is a terminal for a remote user. If the object remotely operated by the user is, for example, a work machine, terminal 3 may be located at a location remote from the work site where the work machine is located, and may be equipped with a display device that displays an image of the work site including the work machine. Note that the remote user is not limited to an operator who remotely operates the work machine as described above. For example, the remote user may be a person involved in the work who uses terminal 3 located at a location remote from the work site where the work machine is located to check the status of work performed by the work machine, monitor the work, manage the work, and guard the work.
[0036] The terminal 3 may be, for example, a mobile terminal such as a handheld terminal, a smartphone, or a tablet terminal, or a wearable terminal such as smart glasses, a head-mounted display, or a smart watch, or may be another type of information terminal. The terminal 3 may be equipped with a photographing device such as a camera that captures images such as color moving images and / or still images.
[0037] The terminal 3 includes an arithmetic processing unit and a memory. Examples of the arithmetic processing unit include a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). Examples of the memory include a RAM (Random Access Memory) and a ROM (Read Only Memory). The terminal 3 includes a control unit. The functions of the control unit are realized by the arithmetic processing unit executing a program stored in the memory. As shown in FIG. 2 , the control unit includes a processing unit, a communication unit, an input unit, an output unit, and a storage unit. The control unit issues execution commands to the processing unit, communication unit, input unit, output unit, storage unit, etc., and the processing unit calculates data and determines the calculation results.
[0038] The terminal 3 includes a communication unit that enables the terminal 3 to communicate with other devices such as the computer 1 and the observation device 2. The communication method may be wireless or wired.
[0039] The terminal 3 has, as an input unit, functions necessary for a remote user to operate the terminal 3. The terminal 3 includes an input device for realizing input. The terminal 3 can be equipped with, for example, an LCD display realizing a touch panel function, a keyboard, a mouse, a pen tablet, hardware buttons on the device, a microphone for voice recognition, etc. The input unit of the terminal 3 is not limited to the input methods described above.
[0040] The terminal 3 has, as an output unit, functions necessary for a remote user to operate the terminal 3. The terminal 3 includes an output device for realizing output. The terminal 3 may have, for example, a display device or an audio output device as the output device. Examples of the display device include a liquid crystal display, a PC display, a projector, etc., and examples of the audio output device include a speaker, etc. The output unit of the terminal 3 is not limited to the output methods described above.
[0041] The terminal 3 includes a storage unit such as a hard disk, a semiconductor memory, a memory card, or other data storage (recording medium). The data may be stored in a cloud service, a database, or the like.
[0042] Furthermore, some or all of the functions of the image-collecting computer 1 may be implemented in the observation device 2 or the terminal 3 by software virtualization.
[0043] The above is the basic concept and basic configuration of the video aggregation system, which is a system including the video aggregation computer 1 (distributed video aggregation computer), the observation device 2, and the terminal 3.
[0044] [Video Aggregation Processing] The video aggregation processing executed by the video aggregation computer 1 of this embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart of the video aggregation processing executed by the video aggregation computer 1.
[0045] The acquisition unit 20 of the video aggregation computer 1 acquires a plurality of videos captured by a plurality of observation devices 2 (step S1). Specifically, the acquisition unit 20 acquires videos of a quality preset by the observation devices 2 from the observation devices 2 via the video channel 4 of the video aggregation computer 1.
[0046] 4, the video aggregation computer 1 includes a video channel 4, which is configured from a plurality of ports 5 (specifically, 11 ports 5), but the number of ports is not limited. The plurality of ports 5 may be set in advance according to a plurality of observation devices 2. In other words, a plurality of observation devices 2 may be associated with a plurality of ports 5. The acquisition unit 20 may acquire the video captured by each observation device 2 at the port 5 corresponding to that observation device 2.
[0047] The setting unit 10 of the video consolidation computer 1 sets the video portion to be displayed on the terminal 3 (e.g., the video portion required by the terminal 3 of the remote user) for each of the acquired videos (step S2). Specifically, as described above, the setting unit 10 may set the video portion based on information from the observation device 2, input provided to the terminal 3 by the user of the terminal 3, input provided to the computer 1 by the user of the computer 1, or an inference result obtained by machine learning. As described above, the video portion may be a specific range (specific area) of the original video. The specific range (the specific area) may be a range (area) that includes a specific object, a range (area) that corresponds to a specific work area, or a range (area) that is set based on other conditions.
[0048] For example, a specific example will be described in which the setting unit 10 sets the video portion based on an input by a user of the terminal 3. In this specific example, as shown in Fig. 4, the computer 1 acquires at least a plurality of videos (specifically, three videos) taken by a plurality of observation devices 2 (specifically, three observation devices 2) that are some of the many observation devices 2 present at the site X, and sets video portions A1, A2, and A3 to be displayed by the terminal 3 for each of the acquired three videos.
[0049] In the above specific example, the user may input to the input unit of terminal 3 to specify a video portion for each of the three videos. This input may include, for example, information specifying a partial range (area) of the original video, or information specifying the entire range (entire area) of the original video. Upon receiving this input, terminal 3 transmits information corresponding to the input (information for specifying the video portion) to video aggregation computer 1, and video aggregation computer 1 receives the transmitted information. That is, the reception unit 22 of computer 1 receives information for specifying the video portion for each of the three videos. Then, based on the received information, setting unit 10 sets the video portion to be displayed on terminal 3 (the video portion required by terminal 3) for each of the three videos.
[0050] It is preferable that at least one of the plurality of video portions (the three video portions A1, A2, and A3) is a part of the original video. Each of the plurality of video portions A1, A2, and A3 may be a part of the original video. Furthermore, one of the plurality of video portions A1, A2, and A3 (e.g., video portion A1) may be the entire original video, and each of the remaining video portions (e.g., video portions A2 and A3) may be a part of the original video.
[0051] In the above specific example, the setting unit 10 sets the video portion based on input by the user of the terminal 3, but even if the setting unit 10 sets the video portion based on input made by the user of the computer 1 to the input unit of the computer 1, the setting unit 10 can set the video portion in the same manner as in the above specific example.
[0052] Next, a brief description will be given of a specific example in which the setting unit 10 sets the video portion based on information from the observation device 2. The computer 1 may detect an object included in the video by performing image recognition on the video acquired from the observation device 2, and set a range including the object (e.g., a range including the object and its surroundings) as the video portion based on the detection results. In this case, the object may be detected based on an inference result using a technique such as machine learning or deep learning, and the setting unit 10 may set a range including the object (e.g., a range including the object and its surroundings) as the video portion based on the detection results.
[0053] When the setting unit 10 sets the video portion, it may also set the type of data (e.g., video only data, video and operation signal data, video and audio data, or video, operation signal, and audio data) included in the video portion to be output to the terminal 3 in the process described below (e.g., step S4 in FIG. 3 ) from the data (video, operation signal, audio, etc.). For example, the setting unit 10 may also set the type of data to be output from the terminal 3 based on conditions such as limitations on the amount of data transmission that can be used, information from the observation device 2, input from a remote user of the terminal 3, inference results by machine learning, etc.
[0054] The synthesis unit 11 of the image integration computer 1 combines the set plurality of image portions into one image (step S3).
[0055] The compositing unit 11 trims (cuts out) at least one of the acquired multiple videos based on settings for that video. The compositing unit 11 may also trim each of the multiple videos based on settings for that video. This allows the compositing unit 11 to generate video data for multiple video portions obtained by trimming the multiple videos. The compositing unit 11 then combines the generated multiple video portions into a single video. The compositing unit 11 may also compress the video data of the combined single video to generate compressed video data for the single video. Furthermore, the compositing unit 11 may compress the video data of each of the generated multiple video portions, and then combine the multiple video portions into a single video.
[0056] For example, when a specific range (specific area) of the original video is set as a setting, the composition unit 11 may generate video data of the video portion by trimming this specific range (specific area) from the original video. The composition unit 11 may generate video data of the video portion by trimming a range of the original video that includes a specific object or a specific work area of the original video from the original video.
[0057] The composition unit 11 may compose multiple video portions generated by trimming into a single video, for example, as follows. For example, the composition unit 11 may set the position, size, etc. of each video portion based on information from the observation device 2, input from a remote user, estimation by machine learning, etc., and combine the multiple video portions into a single video based on the set positions, sizes, etc. When the composition unit 11 combines the multiple video portions, it is preferable to combine the multiple video portions into a single video while the times (timings) at which the acquired multiple videos were taken are synchronized.
[0058] Furthermore, the synthesis unit 11 may store the synthesized video in the video storage unit 30 .
[0059] The synthesizing unit 11 is not limited to performing both trimming and compression, but may be configured to perform only trimming out of the trimming and compression.
[0060] In addition, when combining multiple video portions into a single video, the combining unit 11 may combine the multiple video portions into a single video in a state where the times (timings) at which the original multiple acquired videos were taken are not synchronized.
[0061] The output unit 21 of the video consolidation computer 1 transmits the single video that has been synthesized and encoded, such as compressed, to the terminal 3, and outputs it as a single screen on the terminal 3 (step S4). Specifically, the output unit 21 transmits the single synthesized video to the terminal 3, and the terminal 3 displays the single video on a single screen of the display device of the terminal 3, as in the display example of the terminal 3 in FIG.
[0062] The output unit 21 may extract one synthesized image from the image storage unit 30 and display it on the terminal 3 .
[0063] Furthermore, in an environment where the amount of data transmission available for video transmission is limited, when outputting a single composite video to the terminal 3, the output unit 21 may selectively transmit data included in the single composite video to the terminal 3. For example, the output unit 21 may transmit data related to the video and operation signals included in the single composite video to the terminal 3, and not transmit other data (e.g., data related to audio) to the terminal 3. When the video aggregation computer 1 selectively transmits data to the terminal 3, the output unit 21 may selectively transmit the data, or the combining unit 11 may selectively combine data when combining the single video, and the output unit 21 may transmit the combined video obtained by selectively combining the data. A user, such as a user of the computer 1 or a user of the terminal 3, may input data to the input unit of the computer 1 or the input unit of the terminal 3 to specify the selection of data. In this case, the output unit 21 or the combining unit 11 may select data based on the information specified by the input.
[0064] The remote user does not necessarily have to be one, and may be multiple. In this case, the output unit 21 of the computer 1 may transmit the single composite image to each of multiple terminals 3 used by the multiple remote users, and each terminal 3 may display the single composite image on one screen of its display device. In this case, each terminal 3 may be configured to be able to check information about which terminal 3 used by which remote user the single composite image is displayed on its screen.
[0065] The video aggregation computer 1 or the terminal 3 may receive an input from a remote user to specify a predetermined area of one video being output to the terminal 3, and the output unit 21 of the computer 1 may be configured to temporarily output the original video (video before composition) corresponding to the specified area to the terminal 3. The output unit 21 may be configured to temporarily output the video before composition together with the one composite video, or may be configured to temporarily output only the video before composition. This case will be described below.
[0066] The reception unit 22 of the video aggregation computer 1 receives input from the terminal 3 specifying a specific area of the video that the remote user desires within the synthesized and output video (such as input specifying the range, area, item, area, etc. of the video).
[0067] Based on the received input specifying the predetermined area, the video aggregation computer 1 identifies the original video corresponding to this predetermined area (the original video corresponding to this portion before trimming).
[0068] The output unit 21 of the video consolidation computer 1 transmits the identified original video to the terminal 3, and causes the original video to be temporarily output (displayed) on the terminal 3. Specifically, the output unit 21 transmits the identified original video to the terminal 3, and the terminal 3 displays the original video on the display device of the terminal 3.
[0069] If the output unit 21 determines that there is insufficient communication bandwidth when outputting the original video to the terminal 3 (for example, when the resolution of the original video is too high), it may transmit and output the identified original video to the terminal 3 while temporarily hiding videos other than the identified original video (such as a composite video currently being displayed) on the terminal 3.
[0070] This completes the video aggregation process.
[0071] Next, a specific example of a method in which the compositing unit 11 combines multiple video portions into a single video will be described with reference to Fig. 4. However, the compositing method described below is merely an example, and the method in which the compositing unit 11 combines multiple video portions into a single video is not limited to the specific example below, and other methods may also be employed.
[0072] The screen of terminal 3 in Fig. 4 displays a single image composed of multiple image portions (specifically, three image portions A1, A2, and A3). That is, the single image in Fig. 4 includes a first image portion A1, a second image portion A2, and a third image portion A3.
[0073] The first video portion A1 is composed of a part or all of the first video captured by the first observation device 2. The first video portion A1 is, for example, a part of the first video (original video) that occupies the specific range. The first video (first video data) is composed of a plurality of frames (a plurality of images) arranged in time series. The plurality of frames includes, for example, the first frame to the nth frame arranged in time series. Therefore, the first video portion A1 before synthesis is composed of a plurality of frame portions arranged in time series. The plurality of frame portions includes, for example, the first frame portion to the nth frame portion arranged in time series. Each of the plurality of frame portions is a part or all of the original frame corresponding to that frame portion. Each of the plurality of frame portions is generated, for example, by trimming the original frame corresponding to that frame portion.
[0074] The second video portion A2 is composed of a part or all of the second video captured by the second observation device 2. The second video portion A2 is, for example, a part of the second video (original video) that occupies the specific range. The second video (second video data) is composed of a plurality of frames (a plurality of images) arranged in time series. The plurality of frames includes, for example, the first frame to the nth frame arranged in time series. Therefore, the second video portion A2 before synthesis is composed of a plurality of frame portions arranged in time series. The plurality of frame portions includes, for example, the first frame portion to the nth frame portion arranged in time series. Each of the plurality of frame portions is a part or all of the original frame corresponding to that frame portion. Each of the plurality of frame portions is generated, for example, by trimming the original frame corresponding to that frame portion. The third video portion A3 is similar to the first video portion A1 and the second video portion A2.
[0075] The composition unit 11 may set a plurality of regions in the single image. The composition unit 11 may set a plurality of regions occupied by a plurality of image portions in the single image. In the specific example of FIG. 4 , the plurality of regions includes a first region occupied by a first image portion A1, a second region occupied by a second image portion A2, and a third region occupied by a third image portion A3. In this specific example, the first region is the center region in the single image, the second region is the left region in the single image, and the third region is the right region in the single image. However, the method of dividing the plurality of regions is not limited to the specific example of FIG. 4 . The composition unit 11 may set the plurality of regions based on, for example, an input made by a user, such as a user of the computer 1 or a remote user of the terminal 3, to an input unit of the computer 1 or an input unit of the terminal 3.
[0076] The synthesizing unit 11 generates a first frame of the one video such that a first frame portion of the first video portion A1 before synthesis is located in a first region, a first frame portion of the second video portion A2 before synthesis is located in a second region, and a first frame portion of the third video portion A3 before synthesis is located in a third region. This allows the synthesizing unit 11 to generate the one video (specifically, a part of the one video). The output unit 21 transmits video data of the first frame of the generated one video to the terminal 3.
[0077] Similarly, the synthesizing unit 11 generates a second frame of the one video such that the second frame portion of the first video portion A1 before synthesis is located in the first region, the second frame portion of the second video portion A2 before synthesis is located in the second region, and the second frame portion of the third video portion A3 before synthesis is located in the third region. This allows the synthesizing unit 11 to generate the one video (specifically, a part of the one video). The output unit 21 transmits video data of the second frame of the generated one video to the terminal 3.
[0078] Similarly, the composition unit 11 generates the nth frame of the one video so that the nth frame portion of the first video portion A1 before composition is located in the first region, the nth frame portion of the second video portion A2 before composition is located in the second region, and the nth frame portion of the third video portion A3 before composition is located in the third region. This allows the composition unit 11 to generate the one video (specifically, a part of the one video). The output unit 21 transmits video data of the nth frame of the generated one video to the terminal 3.
[0079] The synthesis unit 11 synthesizes the single video including multiple frames arranged in chronological order from the first frame to the nth frame, for example as described above, and the output unit 21 can sequentially transmit each frame constituting the single video to the terminal 3.
[0080] In the above specific example, the output unit 21 is configured to transmit the video data of the frames that make up the single video one by one to the terminal 3, but it may also be configured to transmit the video data of multiple frames that make up the single video together to the terminal 3, for example.
[0081] Furthermore, when the combining unit 11 combines the three video portions A1, A2, and A3 into the single video, the three frame portions constituting each frame of the single video may be from the same time or from different times. In other words, the three frame portions constituting each frame of the single video may be synchronized or not synchronized.
[0082] The video aggregation computer 1 according to this embodiment synthesizes multiple video segments required by the terminal 3 (remote user) and outputs them as a single screen on the terminal 3. Therefore, this computer 1 can reduce the amount of data transmission of video data sent from the computer 1 to the terminal 3 compared to when multiple videos acquired from multiple observation devices 2 scattered around the site are sent directly from the computer 1 to the terminal 3.
[0083] Furthermore, the image aggregation computer 1 of this embodiment combines multiple image portions set for multiple images from multiple observation devices 2 into a single image, and displays the combined single image on a single screen at the terminal 3. Therefore, even in an environment where there is a limit to the amount of data transmission that can be used for image transmission, even with the minimum necessary infrastructure, the remote user can closely monitor the on-site conditions of the locations, machines, etc. that are displayed in the images that the remote user requires.
[0084] The image aggregation computer 1 of this embodiment stores a single synthesized image and keeps output to the terminal 3 on standby, so that even in a poor communication environment, images can be switched smoothly and without delay, either automatically by the computer 1 or manually by a remote user. Also, by storing image data in the computer 1 (server), it becomes possible to search for images of the subject being photographed.
[0085] The video consolidation computer 1 of this embodiment can instantly access video data that has been stored in the video storage unit 30 and is on standby for output, and output the video related to that video data to the terminal 3. Therefore, even in a site with a poor communication environment, it is possible to smoothly switch videos without delay, either automatically by the computer or manually by a remote user.
[0086] In the image-aggregating computer 1 of this embodiment, when a user of computer 1 or a remote user of terminal 3 wants to switch the image displayed on terminal 3 to another image, the user can input to the input unit of computer 1 or the input unit of terminal 3 to specify the image portion to be displayed on terminal 3 (e.g., the image portion required by terminal 3), and computer 1 can set the image portion based on the information specified by the input, i.e., the information for specifying the image portion. This allows computer 1 to appropriately switch images showing the on-site situation, such as a machine operated by the user or a site monitored by the user.
[0087] The above-described means and functions are realized by a computer 1 including a CPU, memory, various terminals, etc., reading and executing a predetermined program. The program may be provided, for example, in the form of a cloud service provided from one or more terminals via a network, specifically, for example, in the form of SaaS (Software as a Service). The program may also be provided, for example, in the form of a computer-readable recording medium. In this case, the computer 1 may read the program from the recording medium, transfer it to an internal or external recording device, record it, and execute it. The program may also be pre-recorded on a non-transitory recording device (non-transitory recording medium), such as a magnetic disk, optical disk, or magneto-optical disk, and may be configured to be provided to the terminal from the recording device via a communication line.
[0088] However, when a mobile object such as a vehicle communicates with other devices via, for example, a mobile phone network, the available network bandwidth is prone to fluctuation, which can easily cause problems such as packet loss and video delays due to insufficient communication bandwidth. Specifically, when an observation device is mounted on a mobile object or when a user's terminal is mounted on a mobile object, the communication environment is prone to fluctuation. Furthermore, in areas such as mountainous regions and the sea, the communication environment, such as the network communication speed and communication capacity, may not always be favorable. Furthermore, when multiple observation devices, such as multiple cameras, are deployed on-site and video images from the multiple observation devices are transmitted to user terminals via a network, problems such as video delays and degradation of video quality are likely to occur.
[0089] The video aggregation computer 1, video aggregation method, and program according to this embodiment can reduce the amount of data transfer of video data sent from the computer 1 to the terminal 3 when videos from multiple observation devices 2 are displayed on the terminal 3. Therefore, even in an environment where the amount of data transmission available for video is limited, multiple videos can be displayed on the terminal 3 while suppressing problems such as video delays and degradation of video quality.
[0090] Although the embodiments of the present invention have been described above, the present invention is not limited to these embodiments. Furthermore, the effects described in the embodiments of the present invention are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments of the present invention.
[0091] Summary of the embodiment The video aggregation computer 1 having the first feature is a computer for displaying videos from a plurality of distributed observation devices 2 on a terminal 3. This video aggregation computer 1 includes an acquisition unit 20 that acquires a plurality of videos captured by a plurality of observation devices 2, a setting unit 10 that sets a video portion to be displayed on the terminal 3 for each of two or more videos among the acquired plurality of videos, a composition unit 11 that combines the two or more set video portions into a single video, and an output unit 21 that transmits the combined single video to the terminal 3.
[0092] This video aggregation computer 1 sets the video portions to be displayed on terminal 3 for each of two or more videos among the multiple videos acquired from the multiple observation devices 2, synthesizes the set two or more video portions into a single video, and transmits the synthesized single video to terminal 3. Therefore, this computer 1 can reduce the amount of data transmission of video data sent from computer 1 to terminal 3 compared to when multiple videos acquired from the multiple observation devices 2 are sent directly from computer 1 to terminal 3, and can display the single video on a single screen on terminal 3.
[0093] In the video-aggregating computer 1 having the first feature, the setting unit 10 may determine, from the plurality of videos acquired by the acquisition unit 20, two or more videos to be displayed on terminal 3 and videos that should not be displayed on terminal 3. The setting unit 10 may set a video portion to be displayed on terminal 3 for each of the two or more videos to be displayed on terminal 3. The reception unit 22 may receive information for determining, from the plurality of videos, two or more videos to be displayed on terminal 3 and videos that should not be displayed on terminal 3, and the setting unit 10 may set a video portion to be displayed on terminal 3 for each of the two or more videos to be displayed on terminal 3 based on the received information. Specifically, for example, a user such as a user of computer 1 or a user of terminal 3 may make an input to an input unit of computer 1 or an input unit of terminal 3 to specify at least one of the videos to be displayed on terminal 3 and the videos that should not be displayed on terminal 3 from the plurality of videos. Based on the information specified by the input, the setting unit 10 can determine two or more images from the plurality of images to be displayed on the terminal 3 and images that should not be displayed on the terminal 3.
[0094] This video aggregation computer 1 may include an acquisition unit 20 that acquires multiple videos taken by multiple observation devices 2, a setting unit 10 that sets the video portion to be displayed on terminal 3 for each of the acquired multiple videos, a synthesis unit 11 that synthesizes the multiple video portions set for the multiple videos into a single video, and an output unit 21 that transmits the synthesized single video to terminal 3.
[0095] This video aggregation computer 1 sets the video portions to be displayed on terminal 3 for each of the multiple videos acquired from the multiple observation devices 2, synthesizes the multiple video portions set for the multiple videos into a single video, and transmits the single synthesized video to terminal 3. Therefore, this computer 1 can reduce the amount of data transmission of video data sent from computer 1 to terminal 3 compared to when multiple videos acquired from the multiple observation devices 2 are sent directly from computer 1 to terminal 3, and can display the single video on a single screen on terminal 3.
[0096] In the video-intensive computer 1 having the first feature, the output unit 21 may output the one video on one screen of the terminal 3 .
[0097] Specifically, for example, the video aggregation computer 1 may include an acquisition unit 20 that acquires multiple videos taken by multiple observation devices 2, a setting unit 10 that sets the video portions required by the terminal 3 for each of the acquired multiple videos, a synthesis unit 11 that synthesizes the video portions set for each of the multiple videos into a single video, and an output unit 21 that transmits the synthesized single video to the terminal 3 and outputs it as a single screen on the terminal 3.
[0098] This video aggregation computer 1 synthesizes multiple video segments required by terminal 3 (e.g., multiple video segments required by a remote user), transmits the synthesized single video to terminal 3, and causes terminal 3 to output the single video as a single screen. Therefore, this computer 1 can reduce the amount of data transmission of video data transmitted from computer 1 to terminal 3 compared to transmitting multiple videos acquired from multiple observation devices 2 directly from computer 1 to terminal 3. Furthermore, this computer 1 displays, on a single screen on terminal 3, the multiple video segments required by terminal 3 (e.g., multiple video segments required by the user of terminal 3) among the multiple videos from the multiple observation devices 2. As described above, even in an environment with limited data transmission volume available for video transmission, this computer 1 enables detailed monitoring of the on-site conditions of locations, machines, etc., displayed on a single screen on terminal 3 by the multiple video segments required by terminal 3 (e.g., multiple video segments required by a remote user), even with minimal infrastructure.
[0099] In the video aggregation computer 1, the synthesis unit 11 may encode the one video, and the output unit 21 may transmit the encoded one video to the terminal 3. The video aggregation computer 1 may encode the one video (video data) by performing, for example, at least one of compression, encryption, and file format conversion of the one video (video data).
[0100] It is preferable that the synthesis unit 11 of the video aggregation computer 1 synthesizes multiple video portions (two or more video portions) into a single video and then compresses the single video (video data). That is, rather than compressing the video data of each video portion and then synthesizing these video portions into a single video, it is preferable that the synthesis unit 11 synthesizes multiple video portions (two or more video portions) into a single video and then compresses the video data of the single video. The reason for this is as follows: Compression (compression codec) has the property of efficiently compressing portions of video data with small amounts of change and allocating a relatively large bit rate to portions of video data with large amounts of change. In this embodiment, this property is utilized to synthesize multiple video portions into a single video and then compress the single video (video data). This allows for efficient compression (compression codec) according to the distribution of change amounts, thereby further reducing communication load.
[0101] In the video-integrating computer 1 having the second feature, the synthesizing unit 11 may trim at least one of the plurality of acquired videos.
[0102] The video consolidation computer 1 having the second feature trims the original video, leaving only the video portion to be displayed on the terminal 3 (e.g., the video required by the terminal 3 or the remote user) and removing unnecessary portions, thereby reducing the amount of data transmission required for video data sent from the computer 1 to the terminal 3. Therefore, even in an environment where the amount of data transmission available for video transmission is limited, the computer 1 can display, on a single screen on the terminal 3, video of the on-site situation, such as a location or machine, that includes multiple video portions required by the remote user. This allows the user to monitor the on-site situation in detail.
[0103] The image integration computer 1 having the third feature may include an image storage unit 30 for storing the one synthesized image.
[0104] Storing the video portion in the video storage unit 30 may mean storing the video in a standby state in the computer 1, in which the video to be stored can be instantly displayed on the terminal 3 as needed. The standby state is a state in which video is transmitted from the observation device 2 to the computer 1, the computer 1 is not transmitting the video to the terminal 3, and the computer 1 can quickly access the video. The transmission unit 23 can instantly display the video stored in the video storage unit 30 (the video in the standby state) on the terminal 3 as needed.
[0105] With the video consolidation computer 1 having the third feature, for example, the composite video can be stored and the output of the video data can be kept on standby in the computer 1. This allows for smooth video switching without delay, either automatically by the computer 1 or manually by a remote user, even in a location with a poor communication environment. Furthermore, by storing the video data in the computer 1 (e.g., a server), it becomes possible to search for the video of the subject being photographed.
[0106] In the video-intensive computer 1 having the fourth feature, the output unit 21 may transmit the synthesized and stored single video to the terminal 3 and cause the terminal 3 to output the single video on a single screen.
[0107] The video consolidation computer 1 having the fourth feature can instantly access video data that has been stored in the video storage unit 30 and is on standby for output, and output the video related to that video data to the terminal 3. Therefore, even in a site with a poor communication environment, it is possible to smoothly switch videos without delay, either automatically by the computer 1 or manually by a remote user.
[0108] The video aggregation computer 1 having the fifth feature further includes a reception unit 22 that receives information for specifying the video portion, and the setting unit 10 may set the video portion (for example, the video portion required by the terminal 3) based on the received information.
[0109] According to the video-aggregating computer 1 having the fifth feature, it is possible to set the video portion based on the received information. Specifically, for example, when a user of computer 1 or a remote user of terminal 3 wants to switch the video displayed on terminal 3 to another video, the user can input to the input unit of computer 1 or the input unit of terminal 3 to specify the video portion to be displayed on terminal 3, and computer 1 can set the video portion based on the information specified by the input. This allows computer 1 to appropriately switch between videos showing on-site conditions, such as a machine operated by the user or a site monitored by the user.
[0110] In the image-collecting computer 1 having the sixth feature, at least one of the plurality of observation devices 2 may be provided on a moving body.
[0111] In the image aggregation computer 1 having the sixth feature, at least one observation device 2 is mounted on a mobile body, so that the on-site conditions such as locations and machines captured by the observation device 2 mounted on the mobile body can be used as the image portion required by the terminal 3 (for example, the image portion required by a remote user), thereby enabling the user to monitor the on-site conditions.
[0112] A video aggregation method having a seventh feature is a method executed by a computer 1 that displays videos from a plurality of distributed observation devices 2 on a terminal 3. This video aggregation method includes the steps of acquiring a plurality of videos taken by a plurality of observation devices 2, setting video portions to be displayed on the terminal 3 for each of two or more videos among the acquired plurality of videos, combining the two or more set video portions into a single video, and transmitting the combined single video to the terminal 3.
[0113] In this video aggregation method, the computer 1 sets the video portions to be displayed on the terminal 3 for each of two or more videos among the multiple videos acquired from the multiple observation devices 2, combines the two or more set video portions into a single video, and transmits the combined single video to the terminal 3. Therefore, this video aggregation method can reduce the amount of data transmission of video data sent from the computer 1 to the terminal 3 compared to when multiple videos acquired from the multiple observation devices 2 are sent directly from the computer 1 to the terminal 3, and the single video can be displayed on a single screen on the terminal 3.
[0114] This video aggregation method may include the steps of acquiring multiple videos taken by multiple observation devices 2, setting the video portions to be displayed on terminal 3 for each of the acquired multiple videos, combining the multiple video portions set for the multiple videos into a single video, and transmitting the combined single video to terminal 3.
[0115] In this video aggregation method, the computer 1 sets the video portions to be displayed on the terminal 3 for each of the multiple videos acquired from the multiple observation devices 2, synthesizes the multiple video portions set for the multiple videos into a single video, and transmits the single synthesized video to the terminal 3. Therefore, the computer 1 can reduce the amount of data transmission of video data sent from the computer 1 to the terminal 3 compared to sending the multiple videos acquired from the multiple observation devices 2 directly from the computer 1 to the terminal 3, and can display the single video on a single screen on the terminal 3.
[0116] A program having an eighth feature is a program to be executed by a computer 1 to display images from a plurality of distributed observation devices 2 on a terminal 3. This program is a program readable by a computer 1 to cause the computer 1 to execute the steps of acquiring a plurality of images taken by a plurality of observation devices 2, setting an image portion to be displayed on the terminal 3 for each of two or more images among the acquired plurality of images, combining the two or more set image portions into a single image, and transmitting the combined single image to the terminal.
[0117] A computer 1 executing this program sets video portions to be displayed on a terminal 3 for two or more of the multiple videos acquired from multiple observation devices 2, combines the two or more set video portions into a single video, and transmits the combined single video to the terminal 3. Therefore, this program can reduce the amount of data transmission of video data sent from the computer 1 to the terminal 3 compared to when multiple videos acquired from multiple observation devices 2 are directly sent from the computer 1 to the terminal 3, and can display the single video on a single screen on the terminal 3. This program may be recorded in advance on a non-transitory recording device (non-transitory recording medium) such as a magnetic disk, optical disk, or magneto-optical disk, and may be configured to be provided to the terminal from the recording device via a communication line.
[0118] This program may be a program for causing a computer 1 to execute the steps of acquiring multiple images taken by multiple observation devices 2, setting the image portions to be displayed on a terminal 3 for each of the acquired multiple images, combining the multiple image portions set for the multiple images into a single image, and transmitting the combined single image to a terminal 3.
[0119] A computer 1 executing this program sets the image portions to be displayed on a terminal 3 for each of a plurality of images acquired from a plurality of observation devices 2, synthesizes the set image portions for the plurality of images into a single image, and transmits the single synthesized image to the terminal 3. Therefore, this program can reduce the amount of data transmission of image data sent from the computer 1 to the terminal 3 compared to sending a plurality of images acquired from a plurality of observation devices 2 directly from the computer 1 to the terminal 3, and can display the single image on a single screen on the terminal 3.
[0120] The embodiments having the above first to eighth features are in the category of computers, methods, or programs, but similar actions and effects according to the category can also be achieved in other categories such as systems, which will be described later.
[0121] As described above, the computer 1, video aggregation method, and program of this embodiment can reduce the amount of data transmission required for video data sent from the computer 1 to the terminal 3, even in an environment where there is a limit to the amount of data transmission available for video transmission, even with the minimum necessary infrastructure, compared to sending multiple videos acquired from observation devices 2 such as cameras scattered around the site directly from the computer 1 to the terminal 3.
[0122] The present invention further includes the following video aggregation system.
[0123] The video aggregation system comprises a video aggregation computer 1 that acquires multiple images from multiple observation devices 2 that are distributed in a dispersed manner, and a terminal 3 that is located at a distance from the video aggregation computer 1, wherein the video aggregation computer 1 comprises: an acquisition unit 20 that acquires multiple images taken by the multiple observation devices 2; a setting unit 10 that sets the image portions to be displayed on the terminal 3 for each of the acquired multiple images; a synthesis unit 11 that synthesizes the multiple image portions set for the multiple images into a single image; and an output unit 21 that transmits video data for the single synthesized image to the terminal, and the terminal 3 receives the transmitted video data and displays the single image on a single screen using the received video data.
[0124] In this video aggregation system, a video aggregation computer 1 sets the video portions to be displayed on a terminal 3 for each of a plurality of videos acquired from a plurality of observation devices 2, synthesizes the set video portions for the plurality of videos into a single video, and transmits the single synthesized video to a terminal 3. Therefore, this video aggregation system can reduce the amount of data transmission of video data sent from the computer 1 to the terminal 3 compared to when a plurality of videos acquired from a plurality of observation devices 2 are sent directly from the computer 1 to the terminal 3, and can display the single video on a single screen on the terminal 3.
[0125] The video aggregation system may include a plurality of observation devices 2 as components, but the plurality of observation devices 2 are not essential components of the video aggregation system.
[0126] The present invention is not limited to the above-described embodiment, and includes the following modifications, for example.
[0127] [Variation 1] In the video aggregation system according to the embodiment shown in Figures 1 and 4, the site where the observation device 2 is installed is a work site where a work machine 60 such as a shovel, crane, or bulldozer performs work, but this is not limited to the above embodiment. The site where the observation device 2 is installed may also be, for example, a manufacturing site. In this case, a production line for manufacturing products from raw materials is installed at the manufacturing site. In this production line, multiple processes are performed to manufacture products from raw materials. Multiple observation devices 2 are installed at the manufacturing site to capture images of the status of each of the multiple processes.
[0128] Specifically, for example, the production line includes a first process, a second process, and a third process, and the multiple observation devices 2 include a first observation device 2 for photographing the situation of the first process, a second observation device 2 for photographing the situation of the second process, and a third observation device 2 for photographing the situation of the third process.
[0129] In variant example 1, the configurations of the computer 1, the multiple observation devices 2, and the terminal 3 are the same as the configurations of the video aggregation computer 1, the multiple observation devices 2, and the terminal 3 in the embodiment described with reference to Figures 1 to 4, so detailed explanations of these will be omitted.
[0130] In this first modification, similarly to the above embodiment, the image consolidation computer 1 is configured to display images from a plurality of observation devices 2 that are distributed around the manufacturing site on a user's terminal 3. This image consolidation computer 1 includes an acquisition unit 20 that acquires a plurality of images captured by the plurality of observation devices 2, a setting unit 10 that sets the image portions to be displayed on the terminal 3 for each of the acquired images, a composition unit 11 that combines the set image portions for the plurality of images into a single image, and an output unit 21 that transmits the combined image to the terminal 3.
[0131] The terminal 3 is a terminal for a remote user. When multiple observation devices 2 are installed at a manufacturing site, the terminal 3 may be installed at a location away from the manufacturing line and may be equipped with a display device that displays an image of the manufacturing site including the manufacturing line.
[0132] In this first modification, the computer 1 performs the calculation process shown in Fig. 3. Also, in this first modification, the terminal 3 displays, for example, as shown in Figs. 1 and 4 based on instructions from the computer 1.
[0133] Furthermore, the site where the observation device 2 is installed may be, for example, a site where monitoring and maintenance of infrastructure facilities such as a plant or power plant is required.
[0134] [Modification 2] In a video aggregation system according to Modification 2, the computer 1 performs processing to switch at least one of the image quality and size of the video output at the terminal 3 in accordance with at least one of a specific action and a specific video.
[0135] The site where the observation device 2 is placed is a work site, and a work machine 60 such as a shovel, a crane, or a bulldozer is placed at this work site.
[0136] The work machine 60 shown in Fig. 1 is, for example, a shovel. This work machine 60 includes a self-propelled lower traveling body 61, an upper rotating body 62 rotatably supported on the lower traveling body 61, and a working device 63 supported on the upper rotating body 62. In the specific example shown in Fig. 1, the lower traveling body 61 includes a crawler-type traveling device, but the traveling device may also include tires. The working device 63 includes, for example, a boom rotatably supported on the upper rotating body 62, an arm rotatably supported at the tip of the boom, and a bucket rotatably supported at the tip of the arm.
[0137] The work machine 60 is equipped with an attitude detector 64 for detecting the attitude of the work machine 60. The attitude detector 64 may include a boom attitude detector for detecting the attitude of the boom, an arm attitude detector for detecting the attitude of the arm, and a bucket attitude detector for detecting the attitude of the bucket. The attitude detector 64 may further include a rotating unit attitude detector for detecting the attitude of the upper rotating unit 62 relative to the undercarriage 61. The attitude detector 64 may further include an attitude detector that detects the degree of inclination of the work machine 60 relative to a horizontal plane.
[0138] The video aggregation system according to Modification 2 includes a remote control device. The remote control device includes a terminal 3 including a display device as shown in Figures 1 and 4, and a remote controller 65 for remotely operating the work machine 60. The remote controller 65 includes a plurality of remote control levers for moving the work implement 63 of the work machine 60 and rotating the upper rotating body 62. The remote controller 65 also includes at least one of an operation lever and an operation pedal (not shown) for driving the undercarriage 61 of the work machine 60. A user can remotely operate the work machine 60 located at a work site by operating the remote controller 65 of the remote control device while viewing the screen of the display device of the terminal 3.
[0139] The computer 1 according to Modification 2 may perform processing to automatically switch at least one of the image quality and size of the video output on the terminal 3 in accordance with a specific operation of the work machine 60. The specific operation is stored in advance in the computer 1. The user inputs information to the input section of the terminal 3 shown in FIG. 2 to specify the specific operation, and the computer 1 sets the operation corresponding to the input as the specific operation.
[0140] The specific motion may be, for example, at least one of a motion of the working implement 63 and a motion of the upper rotating body 62. Specifically, the specific motion may be a traveling motion of the work machine 60 in a predetermined direction, a traveling motion of the work machine 60 going uphill, or a traveling motion of the work machine 60 going downhill. The specific motion may also be a boom motion, an arm motion, a bucket motion, or a swing motion of the upper rotating body 62. The specific motion may also be a motion in which the traveling speed of the work machine 60 is equal to or greater than a predetermined threshold, a motion in which the operating speed (rotation speed) of the boom, arm, or bucket of the working implement 63 is equal to or greater than a predetermined threshold, or a motion in which the swing speed of the upper rotating body 62 is equal to or greater than a predetermined threshold.
[0141] The computer 1 or the controller of the work machine 60 can determine whether the specific action has been performed based on the detection results input from a motion detection device for determining the specific action. The motion detection device may be, for example, the attitude detector 64, the observation device 2, the remote controller 65, or any other device. If the motion detection device is the observation device 2, the computer 1 or the controller of the work machine 60 may determine whether the specific action has been performed by image recognition of video acquired from the observation device 2. If the motion detection device is the attitude detector 64, the computer 1 or the controller of the work machine 60 can determine whether the specific action has been performed based on the detection results from the attitude detector 64. If the motion detection device is the remote controller 65, the computer 1 or the controller of the work machine 60 can determine whether the specific action has been performed based on the amount of lever operation applied to the remote controller 65.
[0142] When the specific action is performed, the computer 1 may perform processing to automatically change at least one of the image quality and size of the image output by the terminal 3. Specifically, when the specific action is performed, the computer 1 may instruct the terminal 3 to increase the size of the image related to the specific action on the terminal 3 compared to before the specific action was performed, and the terminal 3 may increase the size of the image related to the specific action on the display device of the terminal 3 in accordance with the instruction. When enlarging and displaying the image related to the specific action, the terminal 3 may convert the resolution of the image (original resolution) to a higher resolution using upscaling technology. This can prevent pixelation, block noise, and other problems from occurring in the enlarged and displayed image. This can prevent a decrease in the visibility of the enlarged and displayed image. Note that a known method can be used as the upscaling technology.
[0143] Furthermore, when the specific action is performed, the computer 1 instructs the observation device 2 that captures the video so that the image quality of the video related to the specific action is higher than before the specific action was performed, and the observation device 2 sets the image quality of the video to be higher in accordance with the instruction than before the specific action was performed and provides the video of the set image quality to the computer 1, and the computer 1 uses the video to generate the one video and output the one video to the terminal 3.
[0144] Furthermore, the computer 1 may perform processing to switch at least one of the image quality and size of the image output on the terminal 3 in accordance with a specific image. The computer 1 stores the specific image in advance. The specific image may be, for example, an image whose contrast is equal to or lower than a predetermined threshold, or an image whose resolution is equal to or lower than a predetermined threshold. A user inputs the specific threshold into the input unit of the terminal 3 shown in FIG. 2 , and the computer 1 sets a value corresponding to the input as the specific threshold. The computer 1 or a controller of the terminal 3 can acquire the contrast of the image displayed on the terminal 3.
[0145] When the specific video is displayed on terminal 3, computer 1 may perform processing to automatically switch at least one of the image quality and size of the video output on terminal 3. Specifically, when the specific video is displayed on terminal 3, computer 1 may instruct terminal 3 to increase the size of the specific video on terminal 3 compared to before the specific video was displayed, and terminal 3 may increase the size of the specific video on the display device of terminal 3 in accordance with the instruction compared to before the specific video was displayed. When enlarging the size of the specific video for display, terminal 3 may convert the resolution of the video (original resolution) to a higher resolution using upscaling technology.
[0146] Furthermore, when the specific image is displayed on the terminal 3, the computer 1 may instruct the terminal 3 to increase the contrast of the specific image compared to before the specific image was displayed, and the terminal 3 may increase the contrast of the image in accordance with the instruction compared to before the specific image was displayed.
[0147] [Modification 3] The video aggregation system according to Modification 3 has the following preset function: That is, the computer 1 is configured to be able to save video display settings associated with specific identification conditions.
[0148] The video display setting items may include, for example, a setting item regarding the number of videos (number of video portions) to be displayed on terminal 3, a setting item regarding the size of the videos (size of the video portions) to be displayed on terminal 3, and a setting item regarding the display position of the videos (display position of the video portions) on terminal 3.
[0149] The specific identification condition may be, for example, a condition related to an identifier (identification information) for identifying a user, or a condition related to an identifier (identification information) for identifying a target object such as a work machine, manufacturing equipment, or facility located at a work site. The specific identification condition may also be, for example, a condition related to the operating speed of a remotely controlled target such as the work machine 60, or a condition related to the characteristics of the image. The specific identification condition may also be a condition related to the content of an operation performed at a work site. The content of the operation may include, for example, a work machine operation such as an excavation operation or a reversing operation at a work site, or a vehicle operation such as a transport operation or a reversing operation at a manufacturing site. The specific identification condition may also be a condition related to weather, a condition related to a time period, or a condition related to the quality of the communication environment.
[0150] The user inputs to the input section of the terminal 3 shown in Figure 2 to specify the specific identification condition and the video display setting associated therewith, and the computer 1 associates and saves the specific identification condition with the video display setting based on the input. When the specific identification condition is met, the computer 1 causes the terminal 3 to output video using the video display setting associated with the identification condition. Therefore, the computer 1 can cause the terminal 3 to display video in a display mode that matches the specific identification condition, such as the user's preferences or the type of work machine 60.
[0151] While the above embodiment has been described primarily in terms of the case where the observation device 2 is installed at a work site, environments in which the amount of data transmission is limited may also be, for example, enclosed spaces such as underground spaces, tunnels, and factories, regions in Japan or overseas where communication infrastructure is not adequately developed, space, or regions with congested communications. Furthermore, environments in which the amount of data transmission is limited may also be, for example, environments in which at least one of the observation device 2, computer 1, and terminal 3 communicate while moving.
[0152] The communication means between the computer 1 and the observation device 2 and between the computer 1 and the terminal 3 are not limited to the above-mentioned internet line, a network such as a mobile phone network, or a network constructed by software virtualization. The communication means may be, for example, satellite communication, LPWA (Low Power Wide Area), or other communication means. Satellite communication is a communication method for data communication between the ground and an artificial satellite. LPWA is suitable, for example, when the observation device 2 and the computer 1 are relatively close to each other, or when the computer 1 and the terminal 3 are relatively close to each other, and is expected to have a cost-reducing effect.
[0153] [Variation 5] The video aggregation system according to Variation 5 includes a setting mode and a work mode as control modes. The setting mode is a control mode that is executed before the work mode. The work mode is a control mode that is used when actual work is performed by the work machine 60 at a work site. The computer 1 may set the control mode of the system to the setting mode, and after completing the setting of the video portion (the setting content), switch the control mode of the system from the setting mode to the work mode. Furthermore, the computer 1 may switch the control mode of the system from the work mode to the setting mode when the actual work is completed or when the actual work is interrupted.
[0154] The work mode is a control mode in which the computer 1 generates video data for a video portion of the original video that is to be displayed on the terminal 3 (e.g., step S3 in FIG. 3 ), generates the one video using the video data, transmits the generated one video to the terminal 3 (e.g., step S4 in FIG. 3 ), and displays the one video received by the terminal 3. In this work mode, the computer 1 may transmit to the terminal 3 the video data with a higher image quality than in the setting mode. Also, in this work mode, when a communication line such as a best-effort Internet connection is used, the computer 1 may increase or decrease the image quality of the video data depending on the degree of congestion on the communication line.
[0155] The setting mode is a mode for specifying an image portion. The setting mode is a mode in which the computer 1 can transmit at least one of the multiple images acquired from the multiple observation devices 2 to the terminal 3, the terminal 3 can receive this image, and the received image can be displayed on the display device of the terminal 3. In the setting mode, the computer 1 can display a full-size image (an image including the entire range of the original image) of at least one of the multiple images acquired from the multiple observation devices 2 on the display device of the terminal 3.
[0156] Specifically, for example, when the video data of the video portion is generated by trimming, the setting mode is a mode in which the computer 1 transmits at least one of the multiple videos acquired from the multiple observation devices 2 to the terminal 3 without trimming, the terminal 3 receives this uncropped video, and the received video can be displayed on the display device of the terminal 3. The uncropped video is a full-size video (a video that includes the entire range of the original video).
[0157] In this setting mode, the computer 1 may transmit the full-size video to the terminal 3, for example, with a lower image quality than the image quality of the original video. Also, in this setting mode, the computer 1 may transmit the full-size video to the terminal 3, for example, without changing the image quality of the original video. In other words, the full-size video may be video obtained from the observation device 2 that has been processed to reduce the image quality, or may be video obtained from the observation device 2 that has not been processed to reduce the image quality (video with the same image quality as the original video).
[0158] In this setting mode, terminal 3 accepts input for specifying a video portion to be displayed on the display device of terminal 3 in work mode. Terminal 3 sets the range (area) of the full-size video specified by the accepted input as the video portion (the setting content). Terminal 3 transmits information about the set video portion (the setting content) to computer 1. Computer 1 receives the transmitted information. Computer 1 sets the video portion to be displayed on the display device of terminal 3 in work mode based on the received information. This processing for setting the video portion in setting mode may be performed for each of multiple videos acquired by computer 1 from multiple observation devices 2.
[0159] In this modification 5, the terminal 3 may be provided with a user interface such as a mouse, a keyboard, etc. The user of the terminal 3 may use the user interface to perform the input for specifying a range (area) of the full-size image displayed on the display device of the terminal 3 that the user wishes to set as the image portion (the setting content) within the full-size image displayed on the display device while viewing the full-size image displayed on the display device.
[0160] REFERENCE SIGNS LIST 1 Video aggregation computer 2 Observation device 3 Terminal 4 Video channel 5 Port 10 Setting unit 11 Combining unit 20 Acquisition unit 21 Output unit 22 Reception unit 30 Video storage unit
Claims
1. A video aggregation computer for displaying images from multiple distributed observation devices on a terminal, comprising: an acquisition unit that acquires multiple images taken by the multiple observation devices; a setting unit that sets the image portions to be displayed on the terminal for each of two or more images from the multiple acquired images; a synthesis unit that synthesizes the two or more set image portions into a single image; and an output unit that transmits the synthesized single image to the terminal.
2. The video aggregation computer according to claim 1, wherein said synthesis unit trims at least one of said plurality of acquired videos.
3. The video aggregation computer according to claim 1 or 2, further comprising a video storage unit for storing the one synthesized video.
4. The video aggregation computer according to claim 3, wherein said output unit transmits the synthesized and stored single video to said terminal and causes said terminal to output the single video on a single screen.
5. A video aggregation computer as described in any one of claims 1 to 4, further comprising a reception unit that receives information for specifying the video portion, and the setting unit sets the video portion based on the received information.
6. The image aggregation computer according to any one of claims 1 to 5, wherein at least one of the plurality of observation devices is provided on a moving body.
7. A computer-implemented video aggregation method for displaying video from multiple distributed observation devices on a terminal, the video aggregation method comprising: a step of acquiring multiple videos taken by the multiple observation devices; a step of setting video portions to be displayed on the terminal for each of two or more of the acquired multiple videos; a step of combining the two or more set video portions into a single video; and a step of transmitting the combined single video to the terminal.
8. A computer-readable program for causing a computer to display images from multiple distributed observation devices on a terminal, the computer executing the steps of: acquiring multiple images taken by the multiple observation devices; setting the image portion to be displayed on the terminal for each of two or more images among the acquired multiple images; combining the two or more set image portions into a single image; and transmitting the combined single image to the terminal.
Citation Information
Patent Citations
Image display system
JP2008111269A
Camera network system, display device, and camera
JP2009021959A
Communication system and communication method
JP2009260820A
Image transmission device
JP2012147376A
Video distribution apparatus and method, video distribution system, and video distribution program
JP2014086782A