Remote operation assistance method, remote operation assistance apparatus, and remote operation assistance program

WO2026205591A1PCT designated stage Publication Date: 2026-10-01PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/013375
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-30
Publication Date
2026-10-01

Smart Images

  • Figure JP2026013375_01102026_PF_FP_ABST
    Figure JP2026013375_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A remote operation assistance method is executed by a remote operation assistance system (S) comprising an end device (10) that is mounted on a target object and a control terminal (20) that remotely operates the target object. The end device generates a composite video from at least one camera video obtained by imaging the periphery of the target object by using at least one camera, on the basis of setting information. The end device (10) sends the composite video to the control terminal (20). The control terminal receives the composite video. The control terminal divides the composite video into a plurality of divided videos on the basis of the setting information. An assistance video based on the plurality of divided videos is displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Remote control support method, remote control support apparatus, and remote control support program

[0001] The present disclosure relates to a remote control support method, a remote control support apparatus, and a remote control support program.

[0002] In recent years, transportation solutions using autonomous driving vehicles have attracted attention. For social implementation of such transportation solutions using autonomous driving vehicles, remote monitoring by humans is essential from the perspective of safety. In monitoring operation, videos from a plurality of cameras installed around the periphery of the vehicle to eliminate blind spots are displayed on a control terminal for remote monitoring and operation. An operator appropriately operates / steers the vehicle while watching the video on the terminal.

[0003] Incidentally, in the above monitoring operation, for example, when transmitting video from a vehicle to a control terminal, transmission is performed using a public wireless network. For this reason, it is assumed that there is a possibility that communication quality may not be guaranteed. Accordingly, quality adjustment techniques have been proposed that adjust the video quality to be transmitted (such as resolution, frame rate, and bit rate associated with compression encoding) in accordance with the communication bandwidth, so that the video quality falls within a range that enables monitoring and operation (display delay, image quality, frame rate). Further, in order to support the operator's recognition of the surrounding situation of the vehicle, target information detected by AI and the like are superimposed and displayed on the video.

[0004] Japanese Patent No. 7593806

[0005] For example, in conventional techniques, when videos from a plurality of in-vehicle cameras are individually transmitted from a vehicle to a control terminal, there is a possibility that communication disturbance due to mutual interference of quality adjustment functions, asynchronization due to video transmission delay, or the like may occur.

[0006] An object of the present disclosure is to provide a remote control support method, a remote control support apparatus, and a remote control support program that, when transmitting camera video from a vehicle to a control terminal, can suppress the occurrence of communication disturbance due to mutual interference of quality adjustment functions, asynchronization due to video transmission delay, and the like, as compared with conventional techniques.

[0007] To achieve the above objective, the remote control support method of this disclosure is performed by a remote control support system comprising an end device mounted on a target object and a control terminal for remotely controlling the target object. The end device generates a composite image from at least one camera image captured by at least one camera around the target object based on setting information. The composite image is transmitted to the control terminal. The control terminal receives the composite image. The control terminal divides the composite image into a plurality of segmented images based on the setting information. A support image based on the plurality of segmented images is displayed.

[0008] Figure 1 is a diagram showing an example of the schematic configuration of a remote control support system according to the embodiment. Figure 2 is a diagram showing an example of the configuration of each of the multiple devices included in the remote control support system according to the embodiment. Figure 3 is a flowchart showing an example of the flow of video transmission processing performed by the remote control support system according to the embodiment. Figure 4 is a flowchart showing an example of the flow of video synthesis processing performed by the remote control support system according to the embodiment. Figure 5 is a diagram showing an example of setting information acquired in the video synthesis processing according to the embodiment. Figure 6 is a diagram showing an example of a synthesized video generated in the video synthesis processing according to the embodiment. Figure 7 is a diagram for explaining examples of variations in the synthesis results of the synthesized video according to the embodiment. Figure 8 is for explaining other examples of variations in the synthesis results of the synthesized video according to the embodiment. Figure 9 is for explaining other examples of variations in the synthesis results of the synthesized video according to the embodiment. Figure 10 is for explaining other examples of variations in the synthesis results of the synthesized video according to the embodiment. Figure 11 is a flowchart showing an example of the flow of video quality processing performed by the remote control support system according to the embodiment. Figure 12 is a diagram showing an example of setting information generated in the video quality processing performed by the remote control support system according to the embodiment. Figure 13 is a flowchart showing an example of the flow of video splitting processing performed by the video splitting unit of the control terminal according to the embodiment. Figure 14 shows an example of setting information acquired in the video splitting process performed by the remote control support system according to the embodiment. Figure 15 is a diagram illustrating the cropping process performed in the video splitting process according to the embodiment. Figure 16 shows an example of support video displayed on the control terminal using the video after the video splitting process. Figure 17 shows an example of the flow of AI-based object detection processing performed on the end device side. Figure 18 shows an example of object information generated in the object detection process. Figure 19 shows an example of an object recognized on the composite video based on the object information. Figure 20 shows an example of the flow of AI-based step detection processing performed on the end device side. Figure 21 shows an example of step information generated in the step detection process.Figure 22 shows an example of a step recognized on a composite image based on step information. Figure 23 shows an example of the flow of AI-based area detection processing performed on the end device. Figure 24 shows an example of area information generated in the area detection processing. Figure 25 shows an example of an area recognized on a composite image based on area information. Figure 26 shows an example of the flow of AI-based attention-requiring image detection processing performed on the end device. Figure 27 shows an example of attention-requiring image information generated in the attention-requiring image detection processing. Figure 28 shows an example of attention-requiring image information generated in the attention-requiring image detection processing. Figure 29 shows another example of attention-requiring image information generated in the attention-requiring image detection processing. Figure 30 shows another example of attention-requiring image information generated in the attention-requiring image detection processing. Figure 31 shows an example of the flow of AI-based attention-requiring area expansion processing performed on the end device. Figure 32 shows an example of attention-requiring area video information generated in the attention-requiring area expansion processing. Figure 33 shows an example of a required attention area image enlarged on a composite image based on the required attention area enlargement processing. Figure 34 shows an example of a required attention area image enlarged on a composite image based on the required attention area enlargement processing. Figure 35 shows an example of AI information display on a control terminal. Figure 36 shows an example of AI information display highlighting the attention image obtained by AI processing and the reason why attention is required. Figure 37 shows an example of AI information display highlighting the required attention area extracted by AI processing and the reason why attention is required. Figure 38 shows an example of AI information display highlighting the enlarged required attention area image extracted by AI processing. Figure 39 shows another example of the display format of support video on a control terminal. Figure 40 shows an example of display when color code information is included in the support video. Figure 41 is a flowchart showing an example of the flow of AI processing related to Modification 1. Figure 42 shows an example of video information as setting information generated in the AI ​​processing related to Modification 1. Figure 43 shows an example of a pre-generated interaction correspondence table. Figure 44 is a flowchart showing an example of the AI ​​processing flow related to Modification Example 2.Figure 45 shows an example of video information as setting information generated in the AI ​​processing according to Modification Example 1.

[0009] Hereinafter, the remote control support method, remote control support device, and program according to the embodiments of this disclosure will be described in detail with reference to the attached drawings.

[0010] Figure 1 is a diagram showing an example of the schematic configuration of a remote control support system S according to the embodiment. As shown in Figure 1, the remote control support system S comprises a plurality of end devices 10, a control terminal 20, and a server 30. The remote control support system S is used when the control terminal 20 remotely controls the end devices via the server 30 and the network N.

[0011] In Figure 1, multiple end devices 10 and multiple control terminals 20 are indicated by sub-numbers. The number of end devices 10 and control terminals 20 in the remote operation support system S can be arbitrarily changed according to design conditions, etc. In the following description, when no distinction is made between each end device 10 and each control terminal 20, they will simply be referred to as "end device 10" and "control terminal 20," respectively.

[0012] The end device 10 is an information processing device that can communicate with the server 30 and the control terminal 20 via the network N. The end device 10 is installed on target objects such as vehicles, robots, drones, autonomous mobile devices, manipulators, vehicle interiors, building interiors, and building exteriors. The end device 10 is connected to multiple cameras for imaging the area around the target object and acquires camera images from each camera. Using the camera images acquired from the multiple cameras, the end device 10 performs tasks such as generating a composite image, adjusting the quality (compression, etc.), generating setting information, performing AI (Artificial Intelligence) processing on the composite image, and transmitting the composite image and various information.

[0013] For the sake of clarity in the following explanation, the end device 10 will be assumed to be an information processing device mounted on a vehicle, which is the target object.

[0014] The control terminal 20 is an information processing device that can communicate with the server 30 and the end device 10 via the network N. The control terminal 20 receives synthesized video from the end device 10 via the server 30 and performs video splitting, video conversion, video display, etc. In addition, during the video display process, the control terminal 20 performs superimposed display of AI information received from the end device 10 via the server 30.

[0015] Server 30 is an information processing device that can communicate with end devices 10 and control terminals 20 via network N. Server 30 is also called a backend device and transmits the synthesized video and various information received from end devices 10 to control terminals 20, and stores (archives) the synthesized video and various information received from end devices 10.

[0016] The information processing device described above is typically a computer, and includes, for example, a processor, main memory, auxiliary memory, network interface, device interface, and the like.

[0017] Furthermore, the end device 10, server 20, and control terminal 30 do not necessarily have to be a single information processing device. For example, the various calculations performed by the end device 10, server 20, and control terminal 30 as information processing devices may be performed in parallel using one or more processors, or using multiple computers via a network. Alternatively, the various calculations may be distributed to multiple processing cores within a processor and performed in parallel. In addition, some or all of the processing, functions, means, etc. of this disclosure may be performed by a processor, etc., located on a cloud that can communicate with the end device 10, server 20, and control terminal 30 via a network N.

[0018] Furthermore, in this embodiment, some or all of the configurations and functions of each device may be made up of hardware, or they may be realized by information processing of software (programs) executed by a processor such as a CPU or GPU.

[0019] Figure 2 shows an example of the configuration of each of the multiple devices included in the remote control support system according to the embodiment.

[0020] As shown in Figure 2, the end device 10 is connected to multiple cameras 101 and receives camera images from each camera 101 at a predetermined frame rate. In the following explanation, for the sake of detail, we will assume that the cameras 101 are fisheye cameras provided in each of the four directions of the vehicle on which the end device 10 is mounted: forward, reverse, to the right of the forward direction, and to the left of the forward direction. However, there is no particular limit to the number of cameras 101; there may be one camera or multiple cameras 101.

[0021] The end device 10 includes a video acquisition unit 102, a video synthesis unit 103, a video quality adjustment unit 104, a video transmission unit 105, an AI processing unit 106, an AI information transmission unit 107, and a video synthesis setting transmission unit 108.

[0022] The video acquisition unit 102 receives camera video from each camera 101 at a predetermined frame rate. The video acquisition unit 102 outputs the acquired camera video to the video synthesis unit 103 along with supplementary information such as the camera ID and frame number.

[0023] The video synthesis unit 103 generates a composite image from at least one camera image captured by the camera 101, based on the setting information.

[0024] Here, a composite image refers to a single image created by combining images from multiple cameras. The multiple camera images used to generate the composite image are typically images captured by different cameras, but they may also be multiple images extracted from an image captured by a single camera. In addition to the multiple camera images, the composite image may also include additional information, which will be described later.

[0025] Furthermore, the configuration information sets the information necessary for generating the composite image. For example, the configuration information includes camera configuration information that sets the cropping range for each camera image, video compositing configuration information that sets which camera image and which frame to place in which composite image, scaling configuration information that sets information related to scaling the composite image, and AI configuration information related to the AI ​​information obtained by the AI ​​processing described later. For example, the size and layout of each camera image and additional information on the composite image can be defined by the configuration information.

[0026] The video compositing unit 103 performs a cropping process based on the camera settings information to generate multiple first cropped images from at least one camera image. The video compositing unit 103 may also perform at least one of a fisheye expansion process and a distortion correction process on the multiple first cropped images. The video compositing unit 103 generates a composite image using the multiple first cropped images. Based on the settings information, the video compositing unit 103 generates a composite image by including additional information in the blank areas where no multiple first cropped images exist.

[0027] The video quality adjustment unit 104 performs quality adjustment processing to adjust the resolution, bitrate, and frame rate of the composite video based on the communication bandwidth information. Furthermore, the video quality adjustment unit 104 performs partial processing of the quality adjustment process, targeting a portion of the composite video (for example, a specific camera image portion) as needed. Partial processing of the quality adjustment process is beneficial in that, for example, it can improve communication quality by increasing the compression level of the rear camera image during forward driving, which does not require constant attention.

[0028] The video transmission unit 105 transmits the synthesized video, which has undergone quality adjustment, to the control terminal. The video transmission unit 105 can also add additional information to the video packet during video transmission to synchronize the video with the AI ​​information described later.

[0029] The AI ​​processing unit 106 generates AI information by performing information processing using an AI model (hereinafter also referred to as "AI processing"). Specifically, the AI ​​processing unit 106 uses the synthesized image before quality adjustment processing and the AI ​​model to perform the following processes: object detection processing, step detection processing, area detection processing, attention-grabbing image detection processing, and attention-grabbing area expansion processing. The AI ​​information includes the information obtained by each of these processes.

[0030] Here, object detection processing is the process of detecting the presence and location of objects in the composite image before quality adjustment processing, and then drawing their locations on the composite image. The locations of the detected objects are represented on the composite image, for example, by an emphasized border.

[0031] Furthermore, the step detection process detects the presence and location of uneven objects with a certain height difference or greater on the driving surface (e.g., the ground) shown in the composite image before quality adjustment processing, and then draws their locations on the composite image. Through the step detection process, the locations of the uneven objects are represented on the composite image, for example, as a collection of points.

[0032] Furthermore, region detection processing involves detecting areas in the composite image before quality adjustment processing, classifying them according to their type, generating boundary information (polygons) for each type of region (polygon), and drawing their positions (ranges) on the composite image. The positions of the detected regions are represented on the composite image by methods such as filling each region with a different semi-transparent color according to its type.

[0033] Furthermore, the attention-grabbing image detection process detects the identifier and reason for attention of images that should be focused on from among the images included in the composite image before quality adjustment processing, and highlights those images on the composite image. Through the attention-grabbing image detection process, the images to be focused on and the reasons for attention are highlighted on the composite image, for example, with a thick border.

[0034] Furthermore, the "Attention-Required Area Enlargement Processing" is a process that displays, enlarges, and displays on a separate screen objects, steps, and areas that have been detected as requiring attention, among the objects, steps, and areas detected by the object detection processing, step detection processing, and area detection processing. Through the attention-required area enlargement processing, objects, etc., that have been detected as requiring attention are displayed enlarged on a separate screen in the composite image.

[0035] The AI ​​information transmission unit 107 transmits AI information to the control terminal 20 via the network N and the server 30. If necessary, the AI ​​information can also be transmitted by superimposing it onto the video on the edge device 10 without going through the AI ​​information transmission unit 107, thereby improving the synchronization between the video and the AI ​​information.

[0036] The video synthesis setting transmission unit 108 transmits the setting information used to generate the synthesized video to the control terminal 20 via the network N and the server 30.

[0037] Furthermore, as shown in Figure 2, the server 30 includes a video transmission unit 301, a video storage unit 302, an AI information transmission unit 303, an AI information storage unit 304, a video synthesis setting transmission unit 305, a display setting information transmission unit 306, and a display setting information storage unit 307.

[0038] The video transmission unit 301 transmits video (e.g., composite video) received from the end device 10 via the network N to the control terminal 20.

[0039] The video storage unit 302 stores (archives) video (e.g., composite video) received from the end device 10 via the network N.

[0040] The AI ​​information transmission unit 303 transmits the AI ​​information received from the end device 10 via the network N to the control terminal 20.

[0041] The AI ​​information storage unit 304 stores (archives) the AI ​​information received from the end device 10 via the network N.

[0042] The video synthesis setting transmission unit 305 transmits the video synthesis setting information received from the end device 10 via the network N to the control terminal 20.

[0043] The display setting information transmission unit 306 transmits the display setting information received from the end device 10 via the network N to the control terminal 20.

[0044] The display setting information storage unit 307 stores (archives) the display setting information received from the end device 10 via the network N. However, a user of the system from the outside can intentionally newly input or edit the existing display setting information to be stored.

[0045] Further, as shown in FIG. 2, the control terminal 20 includes a video receiving unit 201, a video dividing unit 202, a video converting unit 203, a video displaying unit 204, an AI information receiving unit 205, an AI information superimposing display unit 206, a display setting unit 207, and a display setting receiving unit 208.

[0046] The video receiving unit 201 receives a video (e.g., a combined video) from the server 30 via the network N.

[0047] The video dividing unit 202 divides the combined video received by the video receiving unit 201 into a plurality of divided videos based on setting information. Further, the video dividing unit 202 generates a plurality of second cut-out videos from the plurality of divided videos based on the setting information. In addition, the video dividing unit 202 executes cut-out processing for cutting out additional information from the combined video received by the video receiving unit 201 based on the setting information.

[0048] The video converting unit 203 executes video conversion processing including at least one of fisheye expansion processing and distortion correction processing on the plurality of second cut-out videos.

[0049] The video display unit 204 displays support videos based on the plurality of divided videos by using the setting information, the second cut-out videos after the video conversion processing, and the additional information obtained by the cut-out processing.

[0050] The AI information receiving unit 205 receives AI information from the server 30 via the network N.

[0051] The AI ​​information overlay display unit 206 draws objects detected by AI processing onto the support image based on the AI ​​information. Examples of specific drawing methods include drawing a frame / rectangle around the location (range) of the detected object, changing the color of the frame / rectangle according to the type of object, and drawing the name of the detected object in text around the frame / rectangle.

[0052] The display setting unit 207 controls the display settings of the support video based on the display setting information received by the display setting receiving unit 208.

[0053] The display setting receiving unit 208 receives display setting information from the server 30 via the network N.

[0054] (Video Transmission Processing) Next, the video transmission processing performed by the remote control support system according to the embodiment will be described. This video transmission processing enables communication with optimized adjustment of the transmission bandwidth without causing mutual interference of the video quality adjustment function, even when transmitting multiple camera images from the end device 10 to the control terminal 20.

[0055] Figure 3 is a flowchart showing an example of the video transmission processing flow performed by the remote control support system according to the embodiment.

[0056] Of the steps shown in Figure 3, steps S1a to S4a and S4b are executed in the end device 10, steps S5a, S5b and S5c are executed in the server 30, and steps S6a, S6b to S10a are executed in the control terminal 20. The specific processing in each step will be described below.

[0057] The video acquisition unit 102 acquires camera images from each camera 101 (step S1a).

[0058] The video synthesis unit 103 performs video synthesis processing using the acquired multiple camera images (step S2a).

[0059] Figure 4 is a flowchart showing an example of the video synthesis process performed by the remote control support system according to the embodiment.

[0060] As shown in Figure 4, the video synthesis unit 103 acquires video from the camera (step S201a).

[0061] The video synthesis unit 103 obtains setting information related to the video from the camera that acquired the video in step S201a (step S202a).

[0062] Figure 5 shows an example of setting information acquired in step S202a of the video synthesis process.

[0063] The first row of Figure 5 shows an example of the setting information acquired in step S202a. As shown in the first row of Figure 5, the setting information acquired in step S202a includes, for example, camera ID, top-left coordinates (horizontal, vertical), bottom-right coordinates (horizontal, vertical), focal length (parameter for fisheye expansion), width after cropping (expanded), and height after cropping (expanded).

[0064] The video synthesis unit 103 performs a cropping process on the video (camera video) acquired in step S201a based on the setting information acquired in step S202a, and generates a plurality of first cropped videos (step S203a).

[0065] The video compositing unit 103 determines whether the processing in steps S201a to S204a has been completed for all cameras (step S205a). If it determines that the processing in steps S201a to S204a has been completed for all cameras (step S205a; Yes), the video compositing unit 103 proceeds to step S206a. On the other hand, if it determines that the processing in steps S201a to S204a has not been completed for all cameras (step S205a; No), the video compositing unit 103 executes each of the processes in steps S201a to S204a for the remaining cameras.

[0066] Furthermore, processing such as fisheye correction and cropping of the display area may be performed so that the user can view it directly. This reduces the amount of data in the composite image and contributes to suppressing transmission bandwidth.

[0067] The series of processes from step S201a to step S205a are repeatedly executed, for example, based on the imaging frame rate of the camera.

[0068] The video synthesis unit 103 acquires pre-set video synthesis setting information (step S206a).

[0069] The second and third rows of Figure 5 show an example of the configuration information acquired in step S206a. As shown in the second and third rows of Figure 5, the configuration information acquired in step S206a includes, for example, information such as the stream ID (the order in which the placement process is executed), the camera ID, and the coordinates of the upper left corner where the video is placed.

[0070] Furthermore, the video compositing settings are dynamically changed depending on the scene (for example, placing a magnified image of a traffic light in the center). Considering the video size, the gaps in the video can be effectively utilized.

[0071] The video synthesis unit 103 acquires the cropped video acquired in step S203a based on the stream ID described in the acquired setting information (step S207a).

[0072] The video synthesis unit 103 places the acquired cropped video into the synthesis area based on the acquired setting information (step S208a).

[0073] The video compositing unit 103 determines whether the processing in steps S206a to S208a has been completed for all camera images (step S209a). If it determines that the processing in steps S206a to S208a has been completed for all camera images (step S209a; Yes), the video compositing unit 103 proceeds to step S210a. On the other hand, if it determines that the processing in steps S206a to S208a has not been completed for all camera images (step S209a; No), the video compositing unit 103 executes the processing in steps S206a to S208a for the remaining camera images.

[0074] The video synthesis unit 103 obtains a synthesized image (step S209a) as a result of performing steps S206a to S208a for each camera image.

[0075] Figure 6 shows an example of a composite image generated in step S210a.

[0076] As a result of performing steps S206a to S208a for each camera image, a composite image is generated in which the first cropped image cut out in step S203a is placed in a predetermined position. For example, if the number of first cropped images is 5, a composite image is generated in which 4 images are arranged vertically and horizontally in a 2-image-per-frame arrangement, as shown in Figure 6, with the remaining 1 image placed in the center of the 4 images.

[0077] There are various variations in how the composite images are displayed. These variations will be explained in detail later.

[0078] The video synthesis unit 103 performs scaling on the generated synthesized video (step S210a).

[0079] The fourth row of Figure 5 shows an example of setting information used in the extension process of step S211a. Based on the setting information shown in the fourth row of Figure 5, the video synthesis unit 103 performs scaling on the synthesized video acquired in step S209a, for example, by interpolation or decimation to an image with a height of 1080 pixels and a width of 1920 pixels.

[0080] The video synthesis unit 103 outputs the synthesized video after the expansion processing to the video transmission unit 105 (step S211a).

[0081] Here, we will explain the variations in the composite image results. Figure 7 is a diagram illustrating examples of variations in the composite image results.

[0082] The upper part of Figure 7 shows an example of a composite result where the images obtained from three or four cameras installed on a vehicle are placed within a rectangle divided into four equal parts.

[0083] The middle section of Figure 7 shows an example of a composite image when, for example, five or six cameras are installed on a vehicle, the images obtained from each camera are placed within a rectangle divided into six equal parts.

[0084] The lower part of Figure 7 shows an example of a composite image when, for example, seven or eight cameras are installed on a vehicle, the images obtained from each camera are placed within a rectangle divided into eight equal parts.

[0085] The composite image shown in Figure 7 is a representation of a circular image obtained using a fisheye lens, placed within a rectangle. As a result, the composite image contains a large amount of blank space in addition to the circular image. This blank space can be used to include various additional information in the composite image.

[0086] Furthermore, for example, fisheye camera images are rectangular data, and outside of the circular area illuminated by external light from the lens, they capture the dark areas inside the camera, resulting in invisible, fine noise. Generally, continuous monochrome images have a high compression ratio, so the amount of information can be reduced by filling in blank areas with a monochrome color.

[0087] Furthermore, the number of camera feeds (the number of cameras) is not particularly limited, as long as the feeds correspond to two or more cameras.

[0088] Figure 8 illustrates another example of the variations in the composite image results.

[0089] In the example of the composite image result in the upper part of Figure 8, additional information is placed in the blank area located in the center, the blank area between adjacent camera images in the vertical direction and the outer edge extending in the vertical direction, and the blank area between adjacent camera images in the horizontal direction and the outer edge extending in the horizontal direction, in the layout of the four camera images shown on the left side of Figure 7.

[0090] Furthermore, in the example of the composite image result in the lower part of Figure 8, additional information is placed in each of the following areas of the layout of the three camera images: the blank area in the center, the blank area between adjacent camera images in the vertical direction and the outer edge extending in the vertical direction, the blank area between adjacent camera images in the horizontal direction and the outer edge extending in the horizontal direction, and the blank area in the lower right where no camera images are placed.

[0091] Furthermore, the areas where additional information can be placed are not limited to blank spaces. For example, some areas of a vehicle captured in camera footage do not contain any information about the surrounding area. Such areas of a vehicle captured in camera footage can also be used as targets for placing additional information.

[0092] Furthermore, any additional information placed in the margins or other areas can be of any kind. For example, it can include color codes in a format that assigns meaning to the shape, arrangement, and time difference when displaying multiple colors, timestamps and video information identifiers in QR codes, AI information (such as the analysis results on the AI ​​side described later), synchronization identifiers, etc.

[0093] Furthermore, the information placed in the margins and other areas is not limited to additional information. For example, if there are five cameras, the image from the fifth camera can be placed in the margin area located in the center, as shown in Figure 6. Figures 9 and 10 illustrate other examples of variations in the composite image results.

[0094] By performing cropping, transformation, and alpha blending of overlapping areas on the camera images acquired by the left, right, front, and rear fisheye cameras shown in Figure 9, it is possible to generate a composite image as a seamless overhead view (combined image) as shown in Figure 10.

[0095] Alternatively, the image can be created by seamlessly combining multiple camera feeds, or by treating the overlapping portions as a single camera feed, or by displaying an overhead view image generated using multiple camera feeds. This allows for the creation of high-quality special effects by processing the uncompressed video.

[0096] Returning to Figure 3, the video quality adjustment unit 104 executes the video quality adjustment process (step S3a).

[0097] Figure 11 is a flowchart showing an example of the video quality processing flow performed by the remote control support system according to the embodiment.

[0098] Figure 12 shows an example of setting information used in video quality processing performed by the remote control support system according to the embodiment. The setting information shown in Figure 12 is pre-stored in a predetermined memory within the end device 10, for example.

[0099] As shown in Figure 11, the video quality adjustment unit 104 acquires communication bandwidth information from the video transmission unit 105, for example, as shown in the first row of Figure 12, which is set to "Maximum transmission amount [Mbps]: 1.5" (step S31a).

[0100] The video quality adjustment unit 104 performs bitrate adjustment processing based on, for example, the bandwidth and bitrate correspondence table shown in the second row of Figure 12 (step S32a).

[0101] The video quality adjustment unit 104 performs a resolution adjustment process based on, for example, the bandwidth and resolution correspondence table shown in the third row of Figure 12 (step S33a).

[0102] The video quality adjustment unit 104 performs frame rate adjustment processing based on, for example, the bandwidth and frame rate correspondence table shown in the second row of Figure 12 (step S34a).

[0103] Returning to Figure 3, the video splitting unit 202 performs the video splitting process (step S7a).

[0104] Figure 13 is a flowchart showing an example of the video splitting process performed by the video splitting unit 202 of the control terminal 20 according to this embodiment.

[0105] As shown in Figure 13, the video splitting unit 202 of the control terminal 20 acquires the composite video from the end device 10 via the server 30 (step S71a).

[0106] The video splitting unit 202 obtains setting information for each camera from the end device 10 via the server 30 (step S72a).

[0107] Figure 14 shows an example of setting information acquired in the video segmentation process performed by the remote control support system according to the embodiment. As shown in the first row of Figure 14, setting information corresponding to each camera (i.e., setting information with the same content as the first row of Figure 5, acquired in step S202a) is acquired on the control terminal 20 side.

[0108] Next, the video splitting unit 202 obtains composite video setting information from the end device 10 via the server 30 (step S73a). That is, as shown in the second and third rows of Figure 14, setting information corresponding to each camera (i.e., setting information with the same content as the second and third rows of Figure 5, obtained in step S202a) is obtained on the control terminal 20 side.

[0109] Next, the video splitting unit 202 performs a process (cutting process) to cut out the display image from the composite image based on the acquired setting information (step S74a).

[0110] Figure 15 is a diagram illustrating the trimming process performed in video splitting.

[0111] As shown in Figure 15, the video splitting unit 202, based on the setting information, for example, as a process corresponding to stream ID "1", performs a cropping process to generate a second cropped video, with the position (0,0) in the composite video being the upper left coordinate of the video corresponding to camera ID "1", and the position (320,240) being the lower right coordinate of the video corresponding to camera ID "1".

[0112] The video splitting unit 202 performs correction processing on the second cropped video acquired by the cropping process based on the setting information (step S75a).

[0113] Returning to Figure 3, the video conversion unit 203 performs video conversion processing on the multiple second cropped images obtained by the video splitting process (step S8a). The conversion process generates each camera image converted to a predetermined size.

[0114] The video display unit 204 displays the support video generated by the video conversion process on the display unit (step S9a).

[0115] Figure 16 shows an example of a support video displayed on a control terminal using video after video segmentation processing.

[0116] As shown in Figure 16, in the support video, each camera image obtained by splitting the composite image through video splitting processing and then converting it through video conversion processing is displayed in a predetermined layout and size. Furthermore, as shown in the upper left of Figure 16, the support video can also include additional information such as a map of the area around the vehicle.

[0117] Furthermore, the video display unit 204 can automatically set the placement of each camera image based on attribute information (front camera image, rear camera image), for example, by displaying the front image in the center. The video display unit 204 can also dynamically change the correction information for the camera images based on vehicle information, AI information, and sensing information, for example, by displaying the rear image larger when the gear is engaged in reverse.

[0118] Alternatively, the placement of each camera image can be automatically determined based on attribute information (front camera image, rear camera image), such as displaying the front image in the center.

[0119] (AI Processing) Next, we will explain the case in which AI processing is performed on the end device 10 side, and the generated AI information is transmitted as video from the end device 10 to the control terminal 20.

[0120] As shown in Figure 3, the AI ​​processing unit 106 executes AI processing (step S3c). Below, we will describe in detail the following typical examples of processing performed by the AI ​​processing unit 106: object detection processing, step detection processing, area detection processing, attention-requiring image detection processing, and attention-requiring area expansion processing.

[0121] Figure 17 shows an example of the flow of AI-based object detection processing performed on the end device side.

[0122] As shown in Figure 17, the AI ​​processing unit 106 acquires the synthesized image from the image synthesis unit 103 (step S301c).

[0123] The AI ​​processing unit 106 performs object detection processing using the acquired synthesized image and the AI ​​(step S302c).

[0124] The AI ​​processing unit 106 outputs the object information obtained by the object detection process to the AI ​​information transmission unit 107 (step S303c).

[0125] Figure 18 shows an example of object information generated during object detection processing. Figure 19 shows an example of an object recognized on a composite image based on the object information.

[0126] As shown in Figure 18, the object information generated in the object detection process includes information such as the object ID, the stream ID in which the object was detected, the top-left coordinate of the object frame in the camera image in the composite video, the width of the object frame, the height of the object frame, the object type, the actual position of the object, and the relative velocity of the object.

[0127] According to the object information described above, for example, object ID "1" is recognized as a "pedestrian" in the image of camera ID "2" in the upper right of the composite image, corresponding to stream ID "2" where the object was detected. The upper left coordinates of the object frame within the camera image in the composite image are (320, 240), and the width of the object frame is "20", the height of the object frame is "50", and the object type is "pedestrian".

[0128] Figure 20 shows an example of the flow of AI-based step detection processing performed on the end device side.

[0129] As shown in Figure 20, the AI ​​processing unit 106 acquires the synthesized image from the image synthesis unit 103 (step 311c).

[0130] The AI ​​processing unit 106 performs step detection processing using the acquired synthesized image and AI (step S312c).

[0131] The AI ​​processing unit 106 outputs the object information obtained by the step detection process to the AI ​​information transmission unit 107 (step S313c).

[0132] Figure 21 shows an example of step information generated during step detection processing. Figure 22 shows an example of a step recognized on a composite image based on the step information.

[0133] As shown in Figure 21, the object information generated in the step detection process includes information such as the stream ID in which the step was detected and the top-left coordinates of the step's position within the camera image in the composite video.

[0134] According to the above step information, for example, as shown in Figure 22, the step is recognized as being located at the top-left coordinates (340, 240) of the step position within the camera image in the composite image, in the image of camera ID "2" in the upper right of the composite image corresponding to stream ID "2" where the step was detected.

[0135] In the examples in Figures 21 and 22, the step is extracted and displayed as a single point on the composite image. Alternatively, depending on the size and extent of the step, it may be extracted and displayed as a line or region with a certain length.

[0136] Figure 23 shows an example of the flow of AI-based region detection processing performed on the end device side.

[0137] As shown in Figure 23, the AI ​​processing unit 106 acquires the synthesized image from the image synthesis unit 103 (step 321c).

[0138] The AI ​​processing unit 106 performs region detection processing using the acquired synthesized image and the AI ​​(step S322c).

[0139] The AI ​​processing unit 106 outputs the object information obtained by the region detection process to the AI ​​information transmission unit 107 (step S323c).

[0140] Figure 24 shows an example of region information generated in region detection processing. Figure 25 shows an example of a region recognized on the composite image based on the region information.

[0141] As shown in Figure 24, the object information generated in the region detection process includes information such as the stream ID in which the region was detected, the polygon of the region within the camera image in the composite video, and the region type.

[0142] According to the above region information, for example, as shown in Figure 25, the region is recognized as a region of type "road" in the polygon [(320, 270), (100, 540), (900, 540), (660, 270)] within the camera image of the region in the composite image, on the image of camera ID "2" in the upper right of the composite image where the region was detected.

[0143] Figure 26 shows an example of the flow of AI-based video detection processing performed on the end device side.

[0144] As shown in Figure 26, the AI ​​processing unit 106 acquires the synthesized image from the image synthesis unit 103 (step 331c).

[0145] The AI ​​processing unit 106 performs attention-grabbing image detection processing using the acquired synthesized image and the AI ​​(step S332c).

[0146] The AI ​​processing unit 106 outputs the attention-requiring video obtained by the attention-requiring video detection process to the AI ​​information transmission unit 107 (step S333c).

[0147] Figure 27 shows an example of attention-grabbing video information generated in the attention-grabbing video detection process. Figure 28 shows an example of attention-grabbing video recognized on the composite image based on the attention-grabbing video detection process.

[0148] As shown in Figure 27, the video information requiring attention generated in the video detection process includes information such as the stream ID that requires attention and the reason why attention is necessary.

[0149] According to the above-mentioned video information requiring attention, for example, as shown in Figure 28, the camera footage in the upper right corner of the composite video, which corresponds to stream ID "2", is recognized as video requiring attention.

[0150] Note that "approaching vehicle" shown in Figure 27 as a video information requiring attention is just one example of a reason why attention is needed. Other reasons that require attention include approaching pedestrians, approaching bicycles, approaching motorcycles, and approaching fallen objects.

[0151] Furthermore, in the attention-grabbing video detection process, it is also possible to output object information that is a factor in determining whether the video requires attention, in addition to the video itself.

[0152] Figure 29 shows another example of attention-grabbing video information generated in the attention-grabbing video detection process. The attention-grabbing video information shown in Figure 29 includes, in addition to the content of the attention-grabbing video information shown in Figure 27, further information such as the stopping cause object ID. The stopping cause object ID is an ID that indicates an object such as a vehicle, pedestrian, bicycle, motorcycle, or fallen object, and multiple IDs may be included depending on the number of stopping causes that occur.

[0153] Figure 30 shows another example of attention-grabbing video information generated in the attention-grabbing video detection process.

[0154] According to the above-mentioned video information requiring attention, for example, as shown in Figure 30, the upper right camera image corresponding to stream ID "2" among the camera images included in the composite image is recognized as video requiring attention. Furthermore, in the video requiring attention, a pedestrian is recognized as stopping factor object ID "5," and an object representing the pedestrian is superimposed.

[0155] Figure 31 shows an example of the flow of AI-driven object magnification processing performed on the end device side.

[0156] As shown in Figure 31, the AI ​​processing unit 106 acquires the synthesized image from the image synthesis unit 103 (step 341c).

[0157] The AI ​​processing unit 106 calculates information on objects requiring attention using the acquired synthesized image and AI (step S342c).

[0158] The AI ​​processing unit 106 outputs the object requiring attention information to the AI ​​information transmission unit 107 (step S343c).

[0159] The AI ​​processing unit 106 extracts the image of the object to be watched from the composite image based on the information of the object to be watched (step S344c).

[0160] Figure 32 shows an example of attention-grabbing area video information generated during the attention-grabbing area expansion process. Figures 33 and 34 show an example of attention-grabbing area video that is expanded on the composite video based on the attention-grabbing area expansion process.

[0161] As shown in Figure 32, the video information of the area requiring attention, generated during the area requiring attention expansion process, includes information such as the enlarged video stream ID in addition to the object information shown in Figure 18.

[0162] According to the above-mentioned video information for the area requiring attention, for example, object ID "1" is recognized as a "pedestrian" in the video of camera ID "2" in the upper right of the composite video corresponding to stream ID "2" where the object was detected, based on the upper-left coordinates (320, 240) of the object frame within the camera image in the composite video, with an object frame width of "20", an object frame height of "50", and an object type of "pedestrian".

[0163] Furthermore, according to the enlarged video stream ID "5" of the video information of the area requiring attention, the video of the area requiring attention is enlarged and displayed in the center of the composite video corresponding to stream ID "5," as shown in Figure 33.

[0164] Furthermore, it is preferable that the area requiring attention be enlarged only when the vehicle's condition or the surrounding environment meets specific conditions. For example, the traffic signal should be enlarged when the vehicle approaches within a certain distance of the pedestrian crossing it is about to cross.

[0165] The attention-grabbing area expansion process allows for the enlargement and display of areas requiring attention, such as pedestrians with white canes and traffic lights. Furthermore, because the image is enlarged before compression, it can generate an enlarged image that is easy to view during decoding.

[0166] Furthermore, if multiple areas requiring attention exist, the area requiring attention enlargement process generates video information for each object ID and performs enlargement display according to each area requiring attention video information.

[0167] For example, Figure 34 shows an example of a composite image display using video information of a region requiring attention for object ID "1" with enlarged video stream ID "5", and video information of a region requiring attention for object ID "2" with enlarged video stream ID "2".

[0168] Next, we will explain the variations in the display format of the AI ​​information obtained through AI processing on the control terminal 20.

[0169] Figure 35 shows an example of AI information display on the control terminal 20. The left side of Figure 35 shows an example where objects detected by object detection processing are highlighted with a colored frame. Operators remotely controlling the control terminal 20 can more easily notice the presence of people or vehicles by observing the camera footage in which people and vehicles are highlighted with colored frames.

[0170] The center of Figure 35 shows an example where steps detected by the step detection process are displayed as dots. Operators remotely controlling the vehicle from the control terminal 20 can easily recognize steps by observing the camera footage in which the step locations are highlighted as dots. As a result, the risk of the vehicle running over steps or other obstacles can be reduced during remote operation.

[0171] The right side of Figure 35 shows an example where the area detected by the area detection process (in Figure 35, the area where the vehicle can travel) is displayed in a different color (hatched in Figure 35). The operator remotely controlling the vehicle from the control terminal 20 can, for example, easily recognize the area where the vehicle can travel by observing the camera image with the area highlighted. As a result, the risk of moving the vehicle outside the drivable area during remote operation can be reduced.

[0172] Figure 36 shows an example of AI information display, highlighting the gazed-on video obtained through AI processing and the reason why attention is needed. As shown in Figure 36, the gazed-on video is highlighted in the center of the lower part of the screen, and the reason why attention is needed, "vehicle approaching," is highlighted in the center of the upper part of the screen. The operator remotely controlling the vehicle from the control terminal 20 can quickly determine which video to look at and why attention is needed on the screen where the support videos are displayed. As a result, the risk of vehicle error can be reduced.

[0173] Figure 37 shows an example of AI information display, highlighting the areas requiring attention extracted by AI processing and the reasons why attention is needed. As shown in Figure 37, the attention-grabbing video and the areas requiring attention within it are highlighted in the center of the lower part of the screen, and the reason for needing attention, "approaching pedestrian," is highlighted in the center of the upper part of the screen. The operator remotely controlling the vehicle from the control terminal 20 can quickly determine which video to view and why attention is needed on the screen displaying the support video. As a result, the risk of vehicle error can be reduced.

[0174] Figure 38 shows an example of AI information display where the enlarged area requiring attention, extracted by AI processing, is highlighted. As shown in Figure 37, the enlarged area requiring attention is highlighted at a predetermined position on the camera image (upper left in Figure 38). The operator remotely controlling the vehicle from the control terminal 20 can quickly determine the presence of people or vehicles by viewing the highlighted enlarged area requiring attention. As a result, the risk of vehicle misoperation and other errors can be reduced.

[0175] Furthermore, if necessary, objects, areas, etc., detected by AI processing and displayed as AI information can be linked to objects, areas, etc., detected by the vehicle's autonomous driving system. This allows the remote operator to see how the autonomous driving system is detecting objects in the surrounding area.

[0176] Figure 39 shows another example of how support video is displayed on a control terminal.

[0177] As shown in Figure 39, the support video displayed on the control terminal 20 may also include the global path of the autonomous vehicle on which the edge device 10 is installed. Including the global path in the support video allows the remote operator to quickly determine where to direct the vehicle during remote operation.

[0178] Furthermore, the support video displayed on the control terminal 20 may also include the local path of the autonomous vehicle equipped with the edge device 10. Including the local path in the support video allows the remote operator to quickly determine how the vehicle will move next during remote operation.

[0179] Furthermore, as another display format for the support video, the same (single) camera footage can be displayed as multiple camera footage by changing the cropping position. This type of display format can be achieved by generating a composite image and corresponding setting information on the upstream edge device 10 by treating the same (single) camera footage as multiple camera footage by changing the cropping position.

[0180] Furthermore, as an alternative display format for the support video, the support video can display not only the video itself, but also video delay information based on color codes and predetermined identifiers. Such a display format can be realized by generating a color code that includes time information and identification information as additional information in the upstream edge device 10. Figure 40 shows an example of display when color code information is included in the support video.

[0181] As shown in Figure 40, color code information is placed and displayed on the support video, for example, in the margin area between camera images. At this time, the smallest unit of the color code is defined so that the length of one side and the placement position are multiples of 16, so as to conform to macroblocks, which are the processing units of common compression encoding methods such as MPEG. Furthermore, when representing time information, it is possible to calculate the video transmission delay by analyzing the color code after reception and calculating the difference between the time of video acquisition and the time on the receiving side.

[0182] (Modification 1) Next, the remote operation support system S according to Modification 1 will be described.

[0183] The remote control support system S according to Modification 1 communicates with the end device 10 and the vehicle on which the end device 10 is installed, and interacts with the vehicle by outputting control based on AI information generated by AI processing to the vehicle.

[0184] Figure 41 is a flowchart showing an example of the AI ​​processing flow according to Modification Example 1.

[0185] As shown in Figure 41, the AI ​​processing unit 106 of the end device 10 acquires a composite image from the image synthesis unit 103 (step S351c).

[0186] The AI ​​processing unit 106 uses the acquired synthesized video and AI to perform a video detection process that requires attention (step S352c). The AI ​​processing unit 106 acquires driving route information from the vehicle and generates video information as setting information (step S353c).

[0187] Figure 42 shows an example of video information as setting information generated in the AI ​​processing according to Modification 1. As shown in Figure 42, the video information includes the stream ID "2" which requires attention, the reason for attention "approaching vehicle", and the driving route information "narrow road". Note that the driving route information "narrow road" is just one example; other driving route information examples include "narrow roadsidewalk", "sidewalk", "road shoulder", "pedestrian crossing", and "railroad crossing".

[0188] The AI ​​processing unit 106 requests interaction from the vehicle based on the generated video information and the pre-generated interaction correspondence table (step S354c). From this point onward, interaction between the end device 10 and the vehicle begins.

[0189] Figure 43 shows an example of a pre-generated interaction correspondence table.

[0190] As shown in Figure 43, the interaction correspondence table sets the corresponding "vehicle interaction" for each combination of "reason for needing attention" and "driving route information". The AI ​​processing unit 106 determines the "vehicle interaction" by referring to the combination of "reason for needing attention" and "driving route information" included in the generated video information and the interaction correspondence table shown in Figure 42. Based on the content of the determined "vehicle interaction", the AI ​​processing unit 106 requests the vehicle to perform the interaction.

[0191] In the above explanation using the example in Figure 41, the end device 10 performs attention-grabbing image detection processing as AI processing, acquires driving route information (map information, etc.) from the vehicle, and initiates an interaction request. In contrast, the interaction may be changed depending on the positional relationship / speed relationship between the detected object and the vehicle, for example, by emitting synthesized speech when a pedestrian approaches in front of the vehicle, but ignoring approaches from behind. Furthermore, sensing information from sensors other than video (LiDAR, radar, GPS, ultrasonic sensors, etc.) and vehicle information (vehicle speed, gear, turn signal, steering value, etc.) can also be used as input to the AI ​​model.

[0192] (Modified Version 2) Next, we will describe the remote operation support system S according to Modified Version 2.

[0193] The remote control support system S according to the modified example 2 communicates between the server 30 and the vehicle equipped with the end device 10, and interacts with the server 30 by outputting control to the vehicle based on AI information generated by advanced AI processing on the server 30 side.

[0194] Figure 44 is a flowchart showing an example of the AI ​​processing flow according to Modification 2. Steps S361c to S364c in Figure 44 are processes on the end device 10 side, steps S365c to S367c are processes on the server 30 side, and step S368 is a process in the vehicle.

[0195] As shown in Figure 44, the AI ​​processing unit 106 of the end device 10 acquires a composite image from the image synthesis unit 103 (step S361c).

[0196] The AI ​​processing unit 106 uses the acquired synthesized image and the AI ​​to perform a process to detect images that require attention (step S362c).

[0197] The AI ​​processing unit 106 acquires driving route information from the vehicle and generates video information as setting information (step S363c).

[0198] Figure 45 shows an example of video information as setting information generated in the AI ​​processing according to Modification Example 1. As shown in Figure 45, the video information includes the stream ID "2" which requires attention, the reason for attention "vehicle approaching", and the driving route information "narrow road".

[0199] The AI ​​processing unit 106 transmits the generated video information to the server 30 and requests interaction with the vehicle from the server 30 (step S364d).

[0200] The AI ​​information transmission unit 303 of the server 30 makes a correspondence determination based on the video information received from the AI ​​processing unit 106 of the end device 10 and, for example, a pre-generated interaction correspondence table shown in Figure 43 (step S365c). At this time, the AI ​​information transmission unit 303 acquires sensing information such as video from the vehicle as necessary.

[0201] Based on the judgment result in step S365c, the AI ​​information transmission unit 303 transmits a response instruction to the control terminal 20 operated by the remote operator, the communication equipment of the responding personnel, etc. (step S367c).

[0202] Furthermore, the AI ​​information transmission unit 303 transmits a corresponding instruction to the vehicle's automated driving system based on the judgment result of step S365c (step S367c).

[0203] The vehicle's autonomous driving system responds to the instruction from the AI ​​information transmission unit 303 and executes the instruction (step S368c). From this point onward, interaction between the server 30 and the vehicle begins.

[0204] In step S367c, instructions sent from the server 30 to the vehicle include, for example, "Instruction to resume / stop driving: Scene example / A pedestrian is nearby, but does not appear to be getting any closer to the vehicle," "Instruction to avoid driving (change of local path): Scene example / A parked car is in front of the vehicle, but does not appear to be moving," "Instruction to change route (change of global path): Scene example / The road ahead is completely blocked due to construction, so it is not possible to continue driving on the current route," and "Instruction on drivable area: Scene example / An area where entry is prohibited due to construction is surrounded by traffic cones, but the AI ​​of the autonomous vehicle cannot determine which area is surrounded and should not be entered (server Examples include: "Using advanced AI to determine areas where driving is not permitted and issuing instructions to the vehicle's AI," "Synthesized voice instructions: Example scene / A pedestrian is nearby but is holding a white cane and is unaware of the vehicle's presence," "Specific instructions to a remote operator (indicating what to do): Example scene / The vehicle is outside the ODD (Operational Design Degree) of autonomous driving and cannot continue driving under any circumstances, so a remote operation request is made to the remote operator / A summary of the vehicle's situation is communicated," and "Instructions for on-site response: Example scene / The vehicle cannot continue driving under autonomous driving or remote response, such as a wheel getting stuck in a ditch, so a rescue request is made to on-site personnel / A summary of the vehicle's situation is communicated."

[0205] As described above, the remote control support system S according to this embodiment comprises an end device 10 mounted on a vehicle as the target object, a control terminal 20 for remotely controlling the vehicle, and a server 30 capable of communicating with the end device 10 and the control terminal 20. Based on the setting information, the end device 10 generates a composite image from at least one camera image captured by at least one camera around the vehicle, and transmits the composite image and the setting information to the control terminal 20 via the server 30. The control terminal 20 receives the composite image and the setting information, divides the composite image into a plurality of segmented images based on the setting information, and displays a support image based on the plurality of segmented images.

[0206] Therefore, since it is not necessary to transmit the video feed from each of the multiple cameras 101 separately, there is no risk of mutual interference of the video quality adjustment function, which can occur when each camera's video feed is transmitted individually. As a result, the benefits of the video quality adjustment function can be maximized, and even when transmitting multiple camera videos from the end device 10 to the control terminal 20, communication with optimized adjustment of the transmission bandwidth can be achieved.

[0207] Furthermore, since there is no need to transmit the video feed from each of the multiple cameras 101 separately, no delay occurs between each video transmission. As a result, asynchronous operation between videos displayed on the control terminal 20 can be suppressed, and a stable remote operation support image can be provided to the user.

[0208] In generating a composite image, the end device 10 generates multiple first cropped images from at least one camera image based on the setting information, and generates a composite image using the multiple first cropped images.

[0209] Therefore, unnecessary parts of the camera footage can be deleted, reducing the overall data size of the composite image.

[0210] In generating a composite image, the end device 10 performs at least one of fisheye processing and distortion correction processing on a plurality of first cropped images to generate a composite image.

[0211] Therefore, it is possible to generate composite images that include camera footage that is highly visible to the user.

[0212] In generating a composite image, the end device 10 generates a composite image in which, for example, if the number of first extracted images is 5, 4 images are arranged vertically and horizontally in a 2-image and 2-image configuration, with the remaining 1 image placed in the center of the 4 images.

[0213] Therefore, it is possible to generate composite images that contain a large amount of video information.

[0214] In generating a composite image, the end device 10 generates a composite image by including additional information in the blank areas where no first extracted images exist, based on the setting information.

[0215] Therefore, even if a composite image contains areas where there is no result information (no image) including, for example, fisheye camera footage, the areas without information can be minimized.

[0216] When the end device 10 transmits a composite video, it transmits the composite video to the control terminal 20 after quality adjustment processing, which adjusts at least one of the resolution, bitrate, or frame rate.

[0217] Therefore, the quality adjustment function enables optimized communication by adjusting the increase or decrease of the transmission bandwidth.

[0218] In the video splitting process, the control terminal 20 generates multiple second-extracted video frames from multiple split video frames and displays a support video in which the multiple second-extracted video frames are arranged.

[0219] Therefore, support images can be generated and displayed using video footage from which unnecessary parts have been removed.

[0220] In the video splitting process, the control terminal 20 performs at least one of fisheye processing and distortion correction processing on multiple second extracted video images.

[0221] Therefore, it is possible to generate and display support images, including camera footage that is highly visible to the user.

[0222] The end device 10 performs AI processing using the synthesized video before quality adjustment processing, which adjusts at least one of the resolution, bitrate, or frame rate, and generates AI information. The end device 10 associates the AI ​​information with the synthesized video and transmits it to the control terminal 20.

[0223] Therefore, since AI processing does not need to be performed on the control terminal 20 side, there is no need to transmit high-resolution video for AI processing from the end device 10 to the control terminal 20. As a result, the quality of the transmitted video can be reduced, and congestion of the communication network bandwidth can be suppressed compared to when AI processing is performed on the control terminal 20 side.

[0224] The end device 10 performs AI processing that includes at least one of the following: object detection processing, step detection processing, area detection processing, attention-requiring image detection processing, and attention-requiring area expansion processing.

[0225] Therefore, the control terminal 20 can provide the user with support images that highlight objects, steps, areas, etc., that should be observed.

[0226] (Modification 3) In the above embodiments and modifications, the case in which the remote control support system S includes a server 30 is illustrated. However, by having the end device 10 and the completed terminal 20 each have the functions of the server 30, the end device 30 and the control terminal 20 can communicate directly, and the server 30 can be omitted.

[0227] (Modification 4) In the above embodiments and modifications, the control terminal 20 has been described as performing a splitting process to divide the composite video into multiple split videos using setting information transmitted from the end device 30 along with the composite video. In contrast, the control terminal 20 may store setting information common to the end device 30 in advance, without transmitting the setting information from the end device, and use this to perform the splitting process.

[0228] In this embodiment, if some or all of the configurations and functions of each device are comprised of software information processing, the software that implements them may be stored in a non-temporary storage medium (non-temporary computer-readable medium) and loaded into a computer, thereby realizing them as software information processing. Alternatively, in this embodiment, some or all of the configurations and functions of each device may be realized through collaborative information processing between software and hardware by implementing the software that implements them in an electronic circuit. Furthermore, each device may download the necessary software via a network N.

[0229] In this embodiment, the program for executing some or all of the configurations and functions of each device may be stored in an HDD (hard disk drive). Alternatively, in this embodiment, the program for executing some or all of the configurations and functions of each device may be pre-installed and provided in ROM.

[0230] Furthermore, in this embodiment, the programs for executing some or all of the configurations and functions of each device may be stored in an installable or executable file format on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD (Digital Versatile Disk), or flexible disk (FD), and provided as a computer program product. Alternatively, in this embodiment, the programs for executing some or all of the configurations and functions of each device may be stored on a computer connected to a network such as the Internet, and provided by allowing downloads via the network. Furthermore, in this embodiment, the programs for executing some or all of the configurations and functions of each device may be provided or distributed via a network such as the Internet.

[0231] Although embodiments have been described above, these embodiments are presented as examples only and are not intended to limit the scope of the invention. This novel embodiment can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. This embodiment and its variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents.

[0232] (Note) The above description of embodiments discloses the following technologies: (1) A remote operation support method performed by a remote operation support system comprising an end device mounted on a target object and a control terminal for remotely operating the target object, the remote operation support method comprising: a composite image generation step in which the end device generates a composite image from at least one camera image captured by at least one camera around the target object based on setting information; a transmission step for transmitting the composite image to the control terminal; a reception step in which the control terminal receives the composite image; a division step in which the control terminal divides the composite image into a plurality of divided images based on the setting information; and a display step for displaying a support image based on the plurality of divided images. (2) The remote operation support method according to (1), wherein the composite image generation step generates a plurality of first cropped images from at least one camera image based on the setting information, and generates the composite image using the plurality of first cropped images. (3) The remote operation support method according to (2), wherein the composite image generation step performs at least one of fisheye processing and distortion correction processing on a plurality of first cropped images to generate the composite image. (4) The remote operation support method according to (2), wherein if the number of the plurality of first cropped images is five, the composite image generation step generates a composite image in which four images are arranged in a 2x2 arrangement, with the remaining one image placed in the center of the four images. (5) The remote operation support method according to (4), wherein the composite image generation step generates the composite image by including additional information in blank areas where no plurality of first cropped images exist, based on the setting information. (6) The remote operation support method according to (1), wherein the transmission step transmits the composite image after quality adjustment processing, which adjusts at least one of the resolution, bitrate, and frame rate, to a control terminal. (7) The remote operation support method according to (1), wherein the division step generates a plurality of second cropped images from a plurality of divided images, and the display step displays the support image in which the plurality of second cropped images are arranged.(8) The remote operation support method according to (7), wherein the division step performs at least one of fisheye processing and distortion correction processing on a plurality of second cropped images. (9) The remote operation support method according to any one of (1) to (8), further comprising an AI processing step in which the end device performs AI processing using the composite image before quality adjustment processing to adjust at least one of resolution, bitrate, and frame rate to generate AI information, and the transmission step transmits the AI ​​information and the composite image to the control terminal. (10) The remote operation support method according to (9), wherein the AI ​​processing step performs AI processing including at least one of object detection processing, step detection processing, area detection processing, attention-requiring image detection processing, and attention-requiring area expansion processing. (11) The remote operation support method according to (9), wherein the transmission step associates the AI ​​information and the composite image and transmits them to the control terminal. (12) A remote operation support system comprising an end device mounted on a target object and a control terminal for remotely operating the target object, wherein the remote operation support device used as the end device comprises: a composite image generation unit that generates a composite image from at least one camera image captured by at least one camera around the target object based on setting information; and a transmission unit that transmits the composite image to the control terminal. (13) A remote operation support program that causes a computer used as the end device to execute: a composite image generation step that generates a composite image from at least one camera image captured by at least one camera around the target object based on setting information; and a transmission step that transmits the composite image to the control terminal.

[0233] 10 End device 20 Control terminal 30 Server 101 Camera 102 Video acquisition unit 103 Video synthesis unit 104 Video quality adjustment unit 105 Video transmission unit 106 AI processing unit 107 AI information transmission unit 108 Video synthesis setting transmission unit 201 Video reception unit 202 Video splitting unit 203 Video conversion unit 204 Video display unit 205 AI information reception unit 206 AI information superimposed display unit 207 Display setting unit 208 Display setting reception unit 301 Video transmission unit 302 Video storage unit 303 AI information transmission unit 304 AI information storage unit 305 Video synthesis setting transmission unit 306 Display setting information transmission unit 307 Display setting information storage unit S Remote control support system

Claims

1. A remote operation support method performed by a remote operation support system comprising an end device mounted on a target object and a control terminal for remotely operating the target object, the remote operation support method comprising: a composite image generation step in which the end device generates a composite image from at least one camera image captured by at least one camera around the target object based on setting information; a transmission step in which the composite image is transmitted to the control terminal; a reception step in which the control terminal receives the composite image; a division step in which the control terminal divides the composite image into a plurality of divided images based on the setting information; and a display step in which a support image based on the plurality of divided images is displayed.

2. The remote operation support method according to claim 1, wherein the composite image generation step generates a plurality of first cropped images from at least one camera image based on the setting information, and generates the composite image using the plurality of first cropped images.

3. The remote operation support method according to claim 2, wherein the composite image generation step involves performing at least one of fisheye processing and distortion correction processing on a plurality of first cropped images to generate the composite image.

4. The remote control support method according to claim 2, wherein the composite image generation step generates a composite image in which, when the number of the multiple first extracted images is five, four images are arranged in a 2x2 arrangement, with the remaining one image placed in the center of the four images.

5. The remote operation support method according to claim 4, wherein the composite image generation step generates the composite image by including additional information in blank areas where no multiple first cropped images exist, based on the setting information.

6. The remote operation support method according to claim 1, wherein the transmission step involves transmitting the synthesized video, after quality adjustment processing which adjusts at least one of the resolution, bitrate, and frame rate, to a control terminal.

7. The remote control support method according to claim 1, wherein the division step generates a plurality of second cropped images from a plurality of divided images, and the display step displays the support image in which the plurality of second cropped images are arranged.

8. The remote control support method according to claim 7, wherein the division step performs at least one of fisheye processing and distortion correction processing on a plurality of second cropped images.

9. The remote operation support method according to any one of claims 1 to 8, further comprising an AI processing step of performing AI processing on the end device using the synthesized video before quality adjustment processing which adjusts at least one of the resolution, bitrate, and frame rate, and generating AI information, wherein the transmission step is to transmit the AI ​​information and the synthesized video to the control terminal.

10. The remote operation support method according to claim 9, wherein the AI ​​processing step performs the AI ​​processing which includes at least one of the following: object detection processing, step detection processing, area detection processing, attention-requiring image detection processing, and attention-requiring area expansion processing.

11. The remote operation support method according to claim 9, wherein the transmission step involves associating the AI ​​information with the synthesized image and transmitting it to the control terminal.

12. A remote control support system comprising an end device mounted on a target object and a control terminal for remotely controlling the target object, wherein the remote control support device used as the end device comprises: a composite image generation unit that generates a composite image from at least one camera image captured by at least one camera around the target object based on setting information; and a transmission unit that transmits the composite image to the control terminal.

13. A remote operation support system comprising an end device mounted on a target object and a control terminal for remotely operating the target object, wherein the remote operation support program causes a computer used as the end device to execute: a composite image generation step of generating a composite image from at least one camera image captured by at least one camera around the target object based on setting information; and a transmission step of transmitting the composite image to the control terminal.